Lawrence Hunter

dblp:h/LawrenceHunter · also Lawrence E. Hunter · DBLP profile ↗
← Back
67ranked-venue papers
8as first author
11since 2021 · last 2026
0000-0003-1455-3370ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 49 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 17 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 1 since 2021
YearPublicationVenuePosition
2026 BeLink: Biomedical Entity Linking Meets Generative Re-Ranking
abstract
Despite recent progress, Biomedical Entity Linking (BEL) with large language models (LLMs) remains computationally inefficient and challenging to deploy in practical settings. In this work, we demonstrate that instruction-tuning of open-source generative models can offer an effective solution when applied at the re-ranking stage of the BEL pipeline. We propose a set-wise instruction-tuning formulation that enables fast and accurate candidate selection. Our method demonstrates strong performance on multiple BEL benchmarks, yielding significant improvements in linking accuracy (3%–24%) while reducing inference time compared to the state-of-the-art. We integrate our generative re-ranker into BeLink, a modular, end-to-end system designed for practical real-world BEL applications.
Darya Shlyk, Stefano Montanelli, Lawrence Hunter
SIGIR3
2026 Improving biomedical entity linking with generative relevance feedback
abstract
MOTIVATION: Biomedical Entity Linking (BEL) maps mentions in biomedical text to standardized identifiers, enabling structured data integration and downstream knowledge discovery. However, current BEL systems remain fundamentally constrained by the recall of the initial candidate pool, where suboptimal retrieval limits the overall effectiveness of the normalization pipeline. RESULTS: We present the first systematic evaluation of Generative Relevance Feedback (GRF) for enhancing candidate retrieval in state-of-the-art BEL systems. GRF leverages large language models (LLMs) to enrich the expressiveness of the mention in a zero-shot fashion. We assess GRF's impact under two scenarios-direct linking prediction and candidate generation in cascading normalization pipelines-and analyze its sensitivity to different LLMs, feedback types, and integration strategies. Experiments across eight corpora and four biomedical knowledge bases demonstrate that integrating GRF significantly improves both accuracy and recall, thereby increasing the upper bound on normalization performance. Our findings highlight GRF as an efficient, model-agnostic solution and underscore its potential as a key component for advancing BEL. AVAILABILITY AND IMPLEMENTATION: The code to reproduce our experiments can be found at: https://doi.org/10.5281/zenodo.17853541.
Darya Shlyk, Lawrence Hunter
Bioinform.2
2025 More Experts Than Galaxies: Conditionally-Overlapping Experts with Biologically-Inspired Fixed Routing
abstract
The evolution of biological neural systems has led to both modularity and sparse coding, which enables energy efficiency and robustness across the diversity of tasks in the lifespan. In contrast, standard neural networks rely on dense, non-specialized architectures, where all model parameters are simultaneously updated to learn multiple tasks, leading to interference. Current sparse neural network approaches aim to alleviate this issue but are hindered by limitations such as 1) trainable gating functions that cause representation collapse, 2) disjoint experts that result in redundant computation and slow learning, and 3) reliance on explicit input or task IDs that limit flexibility and scalability. In this paper we propose Conditionally Overlapping Mixture of ExperTs (COMET), a general deep learning method that addresses these challenges by inducing a modular, sparse architecture with an exponential number of overlapping experts. COMET replaces the trainable gating function used in Sparse Mixture of Experts with a fixed, biologically inspired random projection applied to individual input representations. This design causes the degree of expert overlap to depend on input similarity, so that similar inputs tend to share more parameters. This results in faster learning per update step and improved out-of-sample generalization. We demonstrate the effectiveness of COMET on a range of tasks, including image classification, language modeling, and regression, using several popular deep learning architectures.
Sagi Shaier, Francisco Pereira 0001, Katharina von der Wense, Lawrence Hunter, Matt Jones 0002
ICLR4
2025 GRACKLE: an interpretable matrix factorization approach for biomedical representation learning
abstract
MOTIVATION: Disruption in normal gene expression can contribute to the development of diseases and chronic conditions. However, identifying disease-specific gene signatures can be challenging due to the presence of multiple co-occurring conditions and limited sample sizes. Unsupervised representation learning methods, such as matrix decomposition and deep learning, simplify high-dimensional data into understandable patterns, but often do not provide clear biological explanations. Incorporating prior biological knowledge directly can enhance understanding and address small sample sizes. Nevertheless, current models do not jointly consider prior knowledge of molecular interactions and sample labels. RESULTS: We present GRACKLE, a novel nonnegative matrix factorization approach that applies Graph Regularization Across Contextual KnowLedgE. GRACKLE integrates sample similarity and gene similarity matrices based on sample metadata and molecular relationships, respectively. Simulation studies show GRACKLE outperformed other NMF algorithms, especially with increased background noise. GRACKLE effectively stratified breast tumor samples and identified condition-enriched subgroups in individuals with Down syndrome. The model's latent representations aligned with known biological patterns, such as autoimmune conditions and sleep apnea in Down syndrome. GRACKLE's flexibility allows application to various data modalities, offering a robust solution for identifying context-specific molecular mechanisms in biomedical research. AVAILABILITY AND IMPLEMENTATION: GRACKLE is available at: https://github.com/lagillenwater/GRACKLE.
Lucas A. Gillenwater, Lawrence Hunter, James C. Costello
Bioinform.2
2024 Comparing Template-based and Template-free Language Model Probing
abstract
Sagi Shaier, Kevin Bennett, Lawrence Hunter, Katharina von der Wense. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Sagi Shaier, Kevin Bennett, Lawrence Hunter, Katharina von der Wense
EACL (1)3
2024 Desiderata For The Context Use Of Question Answering Systems
abstract
Prior work has uncovered a set of common problems in state-of-the-art context-based question answering (QA) systems: a lack of attention to the context when the latter conflicts with a model's parametric knowledge, little robustness to noise, and a lack of consistency with their answers.However, most prior work focus on one or two of those problems in isolation, which makes it difficult to see trends across them.We aim to close this gap, by first outlining a set of -previously discussed as well as novel -desiderata for QA models.We then survey relevant analysis and methods papers to provide an overview of the state of the field.The second part of our work presents experiments where we evaluate 15 QA systems on 5 datasets according to all desiderata at once.We find many novel trends, including (1) systems that are less susceptible to noise are not necessarily more consistent with their answers when given irrelevant context; (2) most systems that are more susceptible to noise are more likely to correctly answer according to a context that conflicts with their parametric knowledge; and (3) the combination of conflicting knowledge and noise can reduce system performance by up to 96%.As such, our desiderata help increase our understanding of how these models work and reveal potential avenues for improvements.Code and data can be found here: https://github.com/Shaier/ context_usage_desiderata.git.
Sagi Shaier, Lawrence Hunter, Katharina von der Wense
EACL (1)2
2023 Emerging Challenges in Personalized Medicine: Assessing Demographic Effects on Biomedical Question Answering Systems
abstract
Sagi Shaier, Kevin Bennett, Lawrence Hunter, Katharina Kann. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Sagi Shaier, Kevin Bennett, Lawrence Hunter, Katharina Kann
IJCNLP (1)3
2023 Creating an ignorance-base: Exploring known unknowns in the scientific literature
abstract
BACKGROUND: Scientific discovery progresses by exploring new and uncharted territory. More specifically, it advances by a process of transforming unknown unknowns first into known unknowns, and then into knowns. Over the last few decades, researchers have developed many knowledge bases to capture and connect the knowns, which has enabled topic exploration and contextualization of experimental results. But recognizing the unknowns is also critical for finding the most pertinent questions and their answers. Prior work on known unknowns has sought to understand them, annotate them, and automate their identification. However, no knowledge-bases yet exist to capture these unknowns, and little work has focused on how scientists might use them to trace a given topic or experimental result in search of open questions and new avenues for exploration. We show here that a knowledge base of unknowns can be connected to ontologically grounded biomedical knowledge to accelerate research in the field of prenatal nutrition. RESULTS: We present the first ignorance-base, a knowledge-base created by combining classifiers to recognize ignorance statements (statements of missing or incomplete knowledge that imply a goal for knowledge) and biomedical concepts over the prenatal nutrition literature. This knowledge-base places biomedical concepts mentioned in the literature in context with the ignorance statements authors have made about them. Using our system, researchers interested in the topic of vitamin D and prenatal health were able to uncover three new avenues for exploration (immune system, respiratory system, and brain development) by searching for concepts enriched in ignorance statements. These were buried among the many standard enriched concepts. Additionally, we used the ignorance-base to enrich concepts connected to a gene list associated with vitamin D and spontaneous preterm birth and found an emerging topic of study (brain development) in an implied field (neuroscience). The researchers could look to the field of neuroscience for potential answers to the ignorance statements. CONCLUSION: Our goal is to help students, researchers, funders, and publishers better understand the state of our collective scientific ignorance (known unknowns) in order to help accelerate research through the continued illumination of and focus on the known unknowns and their respective goals for scientific knowledge.
Mayla Boguslav, Nourah M. Salem, Elizabeth K. White, Katherine J. Sullivan, Michael Bada, Teri L. Hernandez, Sonia M. Leach, Lawrence Hunter
J. Biomed. Informatics8
2022 Characterizing Patient Representations for Computational Phenotyping
Tiffany Callahan, Adrianne L. Stefanski, Danielle Ostendorf, Jordan M. Wyrwa, Sara J. Deakyne Davies, George Hripcsak, Lawrence Hunter, Michael G. Kahn
AMIA7
2022 BETA: a comprehensive benchmark for computational drug-target prediction
abstract
Internal validation is the most popular evaluation strategy used for drug-target predictive models. The simple random shuffling in the cross-validation, however, is not always ideal to handle large, diverse and copious datasets as it could potentially introduce bias. Hence, these predictive models cannot be comprehensively evaluated to provide insight into their general performance on a variety of use-cases (e.g. permutations of different levels of connectiveness and categories in drug and target space, as well as validations based on different data sources). In this work, we introduce a benchmark, BETA, that aims to address this gap by (i) providing an extensive multipartite network consisting of 0.97 million biomedical concepts and 8.5 million associations, in addition to 62 million drug-drug and protein-protein similarities and (ii) presenting evaluation strategies that reflect seven cases (i.e. general, screening with different connectivity, target and drug screening based on categories, searching for specific drugs and targets and drug repurposing for specific diseases), a total of seven Tests (consisting of 344 Tasks in total) across multiple sampling and validation strategies. Six state-of-the-art methods covering two broad input data types (chemical structure- and gene sequence-based and network-based) were tested across all the developed Tasks. The best-worst performing cases have been analyzed to demonstrate the ability of the proposed benchmark to identify limitations of the tested methods for running over the benchmark tasks. The results highlight BETA as a benchmark in the selection of computational strategies for drug repurposing and target discovery.
Nansu Zong, Ning Li 0045, Andrew Wen, Victoria Ngo, Yue Yu 0012, Ming Huang 0006, Shaika Chowdhury, Chao Jiang 0002, Sunyang Fu, Richard Weinshilboum, Guoqian Jiang, Lawrence Hunter
Briefings Bioinform.12
2021 Concept recognition as a machine translation problem
abstract
BACKGROUND: Automated assignment of specific ontology concepts to mentions in text is a critical task in biomedical natural language processing, and the subject of many open shared tasks. Although the current state of the art involves the use of neural network language models as a post-processing step, the very large number of ontology classes to be recognized and the limited amount of gold-standard training data has impeded the creation of end-to-end systems based entirely on machine learning. Recently, Hailu et al. recast the concept recognition problem as a type of machine translation and demonstrated that sequence-to-sequence machine learning models have the potential to outperform multi-class classification approaches. METHODS: We systematically characterize the factors that contribute to the accuracy and efficiency of several approaches to sequence-to-sequence machine learning through extensive studies of alternative methods and hyperparameter selections. We not only identify the best-performing systems and parameters across a wide variety of ontologies but also provide insights into the widely varying resource requirements and hyperparameter robustness of alternative approaches. Analysis of the strengths and weaknesses of such systems suggest promising avenues for future improvements as well as design choices that can increase computational efficiency with small costs in performance. RESULTS: Bidirectional encoder representations from transformers for biomedical text mining (BioBERT) for span detection along with the open-source toolkit for neural machine translation (OpenNMT) for concept normalization achieve state-of-the-art performance for most ontologies annotated in the CRAFT Corpus. This approach uses substantially fewer computational resources, including hardware, memory, and time than several alternative approaches. CONCLUSIONS: Machine translation is a promising avenue for fully machine-learning-based concept recognition that achieves state-of-the-art results on the CRAFT Corpus, evaluated via a direct comparison to previous results from the 2019 CRAFT shared task. Experiments illuminating the reasons for the surprisingly good performance of sequence-to-sequence methods targeting ontology identifiers suggest that further progress may be possible by mapping to alternative target concept representations. All code and models can be found at: https://github.com/UCDenver-ccp/Concept-Recognition-as-Translation .
Mayla Boguslav, Negacy D. Hailu, Michael Bada, William A. Baumgartner Jr., Lawrence Hunter
BMC Bioinform.5
2018 Three Dimensions of Reproducibility in Natural Language Processing
Kevin Cohen 0001, Jingbo Xia, Pierre Zweigenbaum, Tiffany Callahan, Orin Hargraves, Foster R. Goss, Nancy Ide, Aurélie Névéol, Cyril Grouin, Lawrence Hunter
LREC10
2017 Reproducibility in Biomedical Natural Language Processing
Kevin Cohen 0001, Aurélie Névéol, Jingbo Xia, Negacy D. Hailu, Cyril Grouin, Lawrence Hunter, Pierre Zweigenbaum
AMIA6
2017 A semantic knowledge-base approach to drug-drug interaction discovery
abstract
We propose a novel, semantic-reasoning-based approach to look for potentially adverse drug-drug interactions (DDIs) by using a knowledge-base of biomedical public ontologies and datasets in a semantic graph representation. This approach makes it possible to find previously unknown relations between different biological entities like drugs, proteins and biological processes, and perform inferences on those relations. Finding nodes that represent drugs in this semantic graph, and intersecting pathways between these nodes (e.g. intersecting at a metabolic pathway step described in Reactome [1] data), can yield to novel drug-drug interactions. The resulting pathways not only describe drug-drug interactions reflected in the literature, but also unstudied interactions that could elucidate reported adverse effects.
Ignacio Tripodi, Kevin Cohen 0001, Lawrence Hunter
BIBM3
2017 Coreference annotation and resolution in the Colorado Richly Annotated Full Text (CRAFT) corpus of biomedical journal articles
abstract
BACKGROUND: Coreference resolution is the task of finding strings in text that have the same referent as other strings. Failures of coreference resolution are a common cause of false negatives in information extraction from the scientific literature. In order to better understand the nature of the phenomenon of coreference in biomedical publications and to increase performance on the task, we annotated the Colorado Richly Annotated Full Text (CRAFT) corpus with coreference relations. RESULTS: The corpus was manually annotated with coreference relations, including identity and appositives for all coreferring base noun phrases. The OntoNotes annotation guidelines, with minor adaptations, were used. Interannotator agreement ranges from 0.480 (entity-based CEAF) to 0.858 (Class-B3), depending on the metric that is used to assess it. The resulting corpus adds nearly 30,000 annotations to the previous release of the CRAFT corpus. Differences from related projects include a much broader definition of markables, connection to extensive annotation of several domain-relevant semantic classes, and connection to complete syntactic annotation. Tool performance was benchmarked on the data. A publicly available out-of-the-box, general-domain coreference resolution system achieved an F-measure of 0.14 (B3), while a simple domain-adapted rule-based system achieved an F-measure of 0.42. An ensemble of the two reached F of 0.46. Following the IDENTITY chains in the data would add 106,263 additional named entities in the full 97-paper corpus, for an increase of 76% percent in the semantic classes of the eight ontologies that have been annotated in earlier versions of the CRAFT corpus. CONCLUSIONS: The project produced a large data set for further investigation of coreference and coreference resolution in the scientific literature. The work raised issues in the phenomenon of reference in this domain and genre, and the paper proposes that many mentions that would be considered generic in the general domain are not generic in the biomedical domain due to their referents to specific classes in domain-specific ontologies. The comparison of the performance of a publicly available and well-understood coreference resolution system with a domain-adapted system produced results that are consistent with the notion that the requirements for successful coreference resolution in this genre are quite different from those of the general domain, and also suggest that the baseline performance difference is quite large.
Kevin Cohen 0001, Arrick Lanfranchi, Miji Joo-young Choi, Michael Bada, William A. Baumgartner Jr., Natalya Panteleyeva, Karin Verspoor, Martha Palmer, Lawrence Hunter
BMC Bioinform.9
2015 KaBOB: ontology-based semantic integration of biomedical databases
abstract
BACKGROUND: The ability to query many independent biological databases using a common ontology-based semantic model would facilitate deeper integration and more effective utilization of these diverse and rapidly growing resources. Despite ongoing work moving toward shared data formats and linked identifiers, significant problems persist in semantic data integration in order to establish shared identity and shared meaning across heterogeneous biomedical data sources. RESULTS: We present five processes for semantic data integration that, when applied collectively, solve seven key problems. These processes include making explicit the differences between biomedical concepts and database records, aggregating sets of identifiers denoting the same biomedical concepts across data sources, and using declaratively represented forward-chaining rules to take information that is variably represented in source databases and integrating it into a consistent biomedical representation. We demonstrate these processes and solutions by presenting KaBOB (the Knowledge Base Of Biomedicine), a knowledge base of semantically integrated data from 18 prominent biomedical databases using common representations grounded in Open Biomedical Ontologies. An instance of KaBOB with data about humans and seven major model organisms can be built using on the order of 500 million RDF triples. All source code for building KaBOB is available under an open-source license. CONCLUSIONS: KaBOB is an integrated knowledge base of biomedical data representationally based in prominent, actively maintained Open Biomedical Ontologies, thus enabling queries of the underlying data in terms of biomedical concepts (e.g., genes and gene products, interactions and processes) rather than features of source-specific data schemas or file formats. KaBOB resolves many of the issues that routinely plague biomedical researchers intending to work with data from multiple data sources and provides a platform for ongoing data integration and development and for formal reasoning over a wealth of integrated biomedical data.
Kevin M. Livingston, Michael Bada, William A. Baumgartner Jr., Lawrence Hunter
BMC Bioinform.4
2015 Visual analysis of biological data-knowledge networks
abstract
BACKGROUND: The interpretation of the results from genome-scale experiments is a challenging and important problem in contemporary biomedical research. Biological networks that integrate experimental results with existing knowledge from biomedical databases and published literature can provide a rich resource and powerful basis for hypothesizing about mechanistic explanations for observed gene-phenotype relationships. However, the size and density of such networks often impede their efficient exploration and understanding. RESULTS: We introduce a visual analytics approach that integrates interactive filtering of dense networks based on degree-of-interest functions with attribute-based layouts of the resulting subnetworks. The comparison of multiple subnetworks representing different analysis facets is facilitated through an interactive super-network that integrates brushing-and-linking techniques for highlighting components across networks. An implementation is freely available as a Cytoscape app. CONCLUSIONS: We demonstrate the utility of our approach through two case studies using a dataset that combines clinical data with high-throughput data for studying the effect of β-blocker treatment on heart failure patients. Furthermore, we discuss our team-based iterative design and development process as well as the limitations and generalizability of our approach.
Corinna Vehlow, David P. Kao, Michael R. Bristow, Lawrence Hunter, Daniel Weiskopf, Carsten Görg
BMC Bioinform.4
2014 Ontology Translation: A Case Study on Translating the Gene Ontology from English to German
Negacy D. Hailu, Kevin Cohen 0001, Lawrence Hunter
NLDB3
2014 Large-scale biomedical concept recognition: an evaluation of current automatic annotators and their parameters
abstract
BACKGROUND: Ontological concepts are useful for many different biomedical tasks. Concepts are difficult to recognize in text due to a disconnect between what is captured in an ontology and how the concepts are expressed in text. There are many recognizers for specific ontologies, but a general approach for concept recognition is an open problem. RESULTS: Three dictionary-based systems (MetaMap, NCBO Annotator, and ConceptMapper) are evaluated on eight biomedical ontologies in the Colorado Richly Annotated Full-Text (CRAFT) Corpus. Over 1,000 parameter combinations are examined, and best-performing parameters for each system-ontology pair are presented. CONCLUSIONS: Baselines for concept recognition by three systems on eight biomedical ontologies are established (F-measures range from 0.14-0.83). Out of the three systems we tested, ConceptMapper is generally the best-performing system; it produces the highest F-measure of seven out of eight ontologies. Default parameters are not ideal for most systems on most ontologies; by changing parameters F-measure can be increased by up to 0.4. Not only are best performing parameters presented, but suggestions for choosing the best parameters based on ontology characteristics are presented.
Christopher S. Funk, William A. Baumgartner Jr., Benjamin Garcia, Christophe Roeder, Michael Bada, Kevin Cohen 0001, Lawrence Hunter, Karin Verspoor
BMC Bioinform.7
2013 Chapter 16: Text Mining for Translational Bioinformatics
abstract
Text mining for translational bioinformatics is a new field with tremendous research potential. It is a subfield of biomedical natural language processing that concerns itself directly with the problem of relating basic biomedical research to clinical practice, and vice versa. Applications of text mining fall both into the category of T1 translational research-translating basic science results into new interventions-and T2 translational research, or translational research for public health. Potential use cases include better phenotyping of research subjects, and pharmacogenomic research. A variety of methods for evaluating text mining applications exist, including corpora, structured test suites, and post hoc judging. Two basic principles of linguistic structure are relevant for building text mining applications. One is that linguistic structure consists of multiple levels. The other is that every level of linguistic structure is characterized by ambiguity. There are two basic approaches to text mining: rule-based, also known as knowledge-based; and machine-learning-based, also known as statistical. Many systems are hybrids of the two approaches. Shared tasks have had a strong effect on the direction of the field. Like all translational bioinformatics software, text mining software for translational bioinformatics can be considered health-critical and should be subject to the strictest standards of quality assurance and software testing.
Kevin Cohen 0001, Lawrence Hunter
PLoS Comput. Biol.2
2013 Rocky Mountain Conference on Bioinformatics Celebrates 10 Years
abstract
The annual Rocky Mountain Conference on Bioinformatics, better known just as “Rocky,” celebrated its tenth meeting in 2012. Rocky has a unique approach in bioinformatics, bringing attention to new scientists and working hard to facilitate the creation of new collaborations and relationships. Set in the spectacular scenery of the Roaring Fork Valley, not far from world-renowned Aspen, Colorado, Rocky offers a world-class locale to accompany great science. Although many of the 127 attendees were from the Rocky Mountain region, scientists came from as far away as Egypt and Korea to join the meeting. Rocky runs two-and-a-half days, organized with keynote addresses bracketing lightning talks and ski breaks. The 2012 keynotes included well-known academic stars Chris Mungall and Olga Troyanskaya, senior thinkers from industry Kirk Jordan and Alex Stewart, and rising young informaticians James Costello, Yuval Itan, and Chris Miller. Miller described novel metagenomic methods that he used to characterize microbiome communities in some of the most polluted sites on earth, describing how to increase sensitivity for detection of rare organisms, and characterizing the surprisingly rich ecosystems found in extremely toxic environments. Itan presented a method for building a network of biological (rather than strictly genetic) relationships among all human genes, then applied graph algorithms to discover disease-causing alleles at genomic scale. He also described impressive experimental validations of his predictions. Costello discussed his analysis of the DREAM competition, where dozens of groups analyzed blinded data regarding gene regulatory networks and genetic determinants of drug response. He demonstrated the value of multiple, independent approaches to difficult prediction problems, and inverted the original goal of the competition to provide insights into the differential predictive power of different sorts of data. Mungall described recent advances in biomedical ontology, which increasingly support complex logical inference, providing complementary information to the traditional statistical analysis of genomic data. Troyanskaya presented a plethora of recent results from her lab, including a remarkable informatics approach to deconvolving cell type– and tissue–specific contributions to gene function and regulation. Jordan described the latest innovations in IBM's high performance computing division, and showcased some impressive applications in medicine, physiological modeling, and sequence analysis. Stewart described SomaLogic's impressive technology that quantitatively assays the abundance of more than 1,000 proteins, demonstrated its clinical utility in stable coronary heart disease, and talked about the informatics challenges involved. In addition to these keynote addresses, 46 scientists from around the world presented 7-minute talks on their research. The purpose of the lightning talks is to briefly introduce the work of each scientist, facilitating more detailed conversations during the abundant free time provided. The topics of the lightning talks spanned a huge range of bioinformatics, from an ENCODE analysis shedding light on the plasticity of replication timing dependent on chromatin structure to an ontological analysis of the discourse structure of scientific journal articles. Topics ranged from phylogenetics to drug response in cancer, and the methods described covered an enormous range, including, for example, mechanistic modeling, Bayesian inference, text mining, and phylogeography. Several of the presenters were giving their first public scientific talk. Rocky is proud of its history of giving the first stage to many young investigators who have gone on to impressive careers. Given the large number of speakers, it is relatively easy to get a chance to present relevant work. There was even one undergraduate researcher reporting on his (impressive!) results. The conference environment sets a relaxed environment for graduate students, postdocs, and other early stage researchers to interact with senior scientists. A remarkably large number of Rocky participants report that new interactions and collaborations arose from their attendance at the meeting, and nearly all say the experience exceeded their expectations. Next year's Rocky is already in preparation, and will be held at the Viceroy Hotel in Snowmass, CO, December 12–14, 2013. The later dates should help with the snow coverage, and the scientific agenda promises to once again be outstanding. We hope to see you there.
Lawrence Hunter
PLoS Comput. Biol.1
2012 Concept annotation in the CRAFT corpus
abstract
BACKGROUND: Manually annotated corpora are critical for the training and evaluation of automated methods to identify concepts in biomedical text. RESULTS: This paper presents the concept annotations of the Colorado Richly Annotated Full-Text (CRAFT) Corpus, a collection of 97 full-length, open-access biomedical journal articles that have been annotated both semantically and syntactically to serve as a research resource for the biomedical natural-language-processing (NLP) community. CRAFT identifies all mentions of nearly all concepts from nine prominent biomedical ontologies and terminologies: the Cell Type Ontology, the Chemical Entities of Biological Interest ontology, the NCBI Taxonomy, the Protein Ontology, the Sequence Ontology, the entries of the Entrez Gene database, and the three subontologies of the Gene Ontology. The first public release includes the annotations for 67 of the 97 articles, reserving two sets of 15 articles for future text-mining competitions (after which these too will be released). Concept annotations were created based on a single set of guidelines, which has enabled us to achieve consistently high interannotator agreement. CONCLUSIONS: As the initial 67-article release contains more than 560,000 tokens (and the full set more than 790,000 tokens), our corpus is among the largest gold-standard annotated biomedical corpora. Unlike most others, the journal articles that comprise the corpus are drawn from diverse biomedical disciplines and are marked up in their entirety. Additionally, with a concept-annotation count of nearly 100,000 in the 67-article subset (and more than 140,000 in the full collection), the scale of conceptual markup is also among the largest of comparable corpora. The concept annotations of the CRAFT Corpus have the potential to significantly advance biomedical text mining by providing a high-quality gold standard for NLP systems. The corpus, annotation guidelines, and other associated resources are freely available at http://bionlp-corpora.sourceforge.net/CRAFT/index.shtml.
Michael Bada, Miriam Eckert, Donald Evans, Kristin Garcia, Krista Shipley, Dmitry Sitnikov, William A. Baumgartner Jr., Kevin Cohen 0001, Karin Verspoor, Judith A. Blake, Lawrence Hunter
BMC Bioinform.11
2012 A corpus of full-text journal articles is a robust evaluation tool for revealing differences in performance of biomedical natural language processing tools
abstract
BACKGROUND: We introduce the linguistic annotation of a corpus of 97 full-text biomedical publications, known as the Colorado Richly Annotated Full Text (CRAFT) corpus. We further assess the performance of existing tools for performing sentence splitting, tokenization, syntactic parsing, and named entity recognition on this corpus. RESULTS: Many biomedical natural language processing systems demonstrated large differences between their previously published results and their performance on the CRAFT corpus when tested with the publicly available models or rule sets. Trainable systems differed widely with respect to their ability to build high-performing models based on this data. CONCLUSIONS: The finding that some systems were able to train high-performing models based on this corpus is additional evidence, beyond high inter-annotator agreement, that the quality of the CRAFT corpus is high. The overall poor performance of various systems indicates that considerable work needs to be done to enable natural language processing systems to work well when the input is full-text journal articles. The CRAFT corpus provides a valuable resource to the biomedical natural language processing community for evaluation and training of new models for biomedical full text publications.
Karin Verspoor, Kevin Cohen 0001, Arrick Lanfranchi, Colin Warner, Helen L. Johnson 0001, Christophe Roeder, Jinho D. Choi, Christopher S. Funk, Yuriy Malenkiy, Miriam Eckert, Nianwen Xue, William A. Baumgartner Jr., Michael Bada, Martha Palmer, Lawrence Hunter
BMC Bioinform.15
2012 AMIA Board white paper: definition of biomedical informatics and specification of core competencies for graduate education in the discipline
abstract
The AMIA biomedical informatics (BMI) core competencies have been designed to support and guide graduate education in BMI, the core scientific discipline underlying the breadth of the field's research, practice, and education. The core definition of BMI adopted by AMIA specifies that BMI is 'the interdisciplinary field that studies and pursues the effective uses of biomedical data, information, and knowledge for scientific inquiry, problem solving and decision making, motivated by efforts to improve human health.' Application areas range from bioinformatics to clinical and public health informatics and span the spectrum from the molecular to population levels of health and biomedicine. The shared core informatics competencies of BMI draw on the practical experience of many specific informatics sub-disciplines. The AMIA BMI analysis highlights the central shared set of competencies that should guide curriculum design and that graduate students should be expected to master.
Casimir A. Kulikowski, Edward H. Shortliffe, Leanne M. Currie, Peter L. Elkin, Lawrence Hunter, Todd R. Johnson, Ira J. Kalet, Leslie Lenert, Mark A. Musen, Judy G. Ozbolt, Jack W. Smith, Peter Tarczy-Hornoch, Jeffrey J. Williamson
J. Am. Medical Informatics Assoc.5
2011 U-Compare bio-event meta-service: compatible BioNLP event extraction services
abstract
BACKGROUND: Bio-molecular event extraction from literature is recognized as an important task of bio text mining and, as such, many relevant systems have been developed and made available during the last decade. While such systems provide useful services individually, there is a need for a meta-service to enable comparison and ensemble of such services, offering optimal solutions for various purposes. RESULTS: We have integrated nine event extraction systems in the U-Compare framework, making them intercompatible and interoperable with other U-Compare components. The U-Compare event meta-service provides various meta-level features for comparison and ensemble of multiple event extraction systems. Experimental results show that the performance improvements achieved by the ensemble are significant. CONCLUSIONS: While individual event extraction systems themselves provide useful features for bio text mining, the U-Compare meta-service is expected to improve the accessibility to the individual systems, and to enable meta-level uses over multiple event extraction systems such as comparison and ensemble.
Yoshinobu Kano, Jari Björne, Filip Ginter, Tapio Salakoski, Ekaterina Buyko, Udo Hahn, Kevin Cohen 0001, Karin Verspoor, Christophe Roeder, Lawrence Hunter, Halil Kilicoglu, Sabine Bergler, Sofie Van Landeghem, Thomas Van Parys, Yves Van de Peer, Makoto Miwa, Sophia Ananiadou, Mariana L. Neves, Alberto D. Pascual-Montano, Arzucan Özgür, Dragomir R. Radev, Sebastian Riedel 0001, Rune Sætre, Hong-Woo Chun, Jin-Dong Kim, Sampo Pyysalo, Tomoko Ohta, Jun'ichi Tsujii
BMC Bioinform.10
2011 High-Precision Biological Event Extraction: Effects of System and of Data
abstract
We approached the problems of event detection, argument identification, and negation and speculation detection in the BioNLP'09 information extraction challenge through concept recognition and analysis. Our methodology involved using the OpenDMAP semantic parser with manually written rules. The original OpenDMAP system was updated for this challenge with a broad ontology defined for the events of interest, new linguistic patterns for those events, and specialized coordination handling. We achieved state-of-the-art precision for two of the three tasks, scoring the highest of 24 teams at precision of 71.81 on Task 1 and the highest of 6 teams at precision of 70.97 on Task 2. We provide a detailed analysis of the training data and show that a number of trigger words were ambiguous as to event type, even when their arguments are constrained by semantic class. The data is also shown to have a number of missing annotations. Analysis of a sampling of the comparatively small number of false positives returned by our system shows that major causes of this type of error were failing to recognize second themes in two-theme events, failing to recognize events when they were the arguments to other events, failure to recognize nontheme arguments, and sentence segmentation errors. We show that specifically handling coordination had a small but important impact on the overall performance of the system. The OpenDMAP system and the rule set are available at http://bionlp.sourceforge.net.
Kevin Cohen 0001, Karin Verspoor, Helen L. Johnson 0001, Christophe Roeder, Philip V. Ogren, William A. Baumgartner Jr., Elizabeth K. White, Hannah J. Tipney, Lawrence Hunter
Comput. Intell.9
2011 Desiderata for ontologies to be used in semantic annotation of biomedical documents
Michael Bada, Lawrence Hunter
J. Biomed. Informatics2
2010 Visualization and Language Processing for Supporting Analysis across the Biomedical Literature
Carsten Görg, Hannah J. Tipney, Karin Verspoor, William A. Baumgartner Jr., Kevin Cohen 0001, John T. Stasko, Lawrence Hunter
KES (4)7
2010 Test Suite Design for Biomedical Ontology Concept Recognition Systems
Kevin Cohen 0001, Christophe Roeder, William A. Baumgartner Jr., Lawrence Hunter, Karin Verspoor
LREC4
2010 A UIMA wrapper for the NCBO annotator
abstract
SUMMARY: The Unstructured Information Management Architecture (UIMA) framework and web services are emerging as useful tools for integrating biomedical text mining tools. This note describes our work, which wraps the National Center for Biomedical Ontology (NCBO) Annotator-an ontology-based annotation service-to make it available as a component in UIMA workflows. AVAILABILITY: This wrapper is freely available on the web at http://bionlp-uima.sourceforge.net/ as part of the UIMA tools distribution from the Center for Computational Pharmacology (CCP) at the University of Colorado School of Medicine. It has been implemented in Java for support on Mac OS X, Linux and MS Windows.
Christophe Roeder, Clément Jonquet, Nigam H. Shah, William A. Baumgartner Jr., Karin Verspoor, Lawrence Hunter
Bioinform.6
2010 The structural and content aspects of abstracts versus bodies of full text journal articles are different
abstract
BACKGROUND: An increase in work on the full text of journal articles and the growth of PubMedCentral have the opportunity to create a major paradigm shift in how biomedical text mining is done. However, until now there has been no comprehensive characterization of how the bodies of full text journal articles differ from the abstracts that until now have been the subject of most biomedical text mining research. RESULTS: We examined the structural and linguistic aspects of abstracts and bodies of full text articles, the performance of text mining tools on both, and the distribution of a variety of semantic classes of named entities between them. We found marked structural differences, with longer sentences in the article bodies and much heavier use of parenthesized material in the bodies than in the abstracts. We found content differences with respect to linguistic features. Three out of four of the linguistic features that we examined were statistically significantly differently distributed between the two genres. We also found content differences with respect to the distribution of semantic features. There were significantly different densities per thousand words for three out of four semantic classes, and clear differences in the extent to which they appeared in the two genres. With respect to the performance of text mining tools, we found that a mutation finder performed equally well in both genres, but that a wide variety of gene mention systems performed much worse on article bodies than they did on abstracts. POS tagging was also more accurate in abstracts than in article bodies. CONCLUSIONS: Aspects of structure and content differ markedly between article abstracts and article bodies. A number of these differences may pose problems as the text mining field moves more into the area of processing full-text articles. However, these differences also present a number of opportunities for the extraction of data types, particularly that found in parenthesized text, that is present in article bodies but not in article abstracts.
Kevin Cohen 0001, Helen L. Johnson 0001, Karin Verspoor, Christophe Roeder, Lawrence Hunter
BMC Bioinform.5
2010 Current methodologies for translational bioinformatics
Yves A. Lussier, Atul J. Butte, Lawrence Hunter
J. Biomed. Informatics3
2010 Exploring Species-Based Strategies for Gene Normalization
abstract
We introduce a system developed for the BioCreative II.5 community evaluation of information extraction of proteins and protein interactions. The paper focuses primarily on the gene normalization task of recognizing protein mentions in text and mapping them to the appropriate database identifiers based on contextual clues. We outline a ""fuzzy" dictionary lookup approach to protein mention detection that matches regularized text to similarly regularized dictionary entries. We describe several different strategies for gene normalization that focus on species or organism mentions in the text, both globally throughout the document and locally in the immediate vicinity of a protein mention, and present the results of experimentation with a series of system variations that explore the effectiveness of the various normalization strategies, as well as the role of external knowledge sources. While our system was neither the best nor the worst performing system in the evaluation, the gene normalization strategies show promise and the system affords the opportunity to explore some of the variables affecting performance on the BCII.5 tasks.
Karin Verspoor, Christophe Roeder, Helen L. Johnson 0001, Kevin Cohen 0001, William A. Baumgartner Jr., Lawrence Hunter
IEEE ACM Trans. Comput. Biol. Bioinform.6
2009 U-Compare: share and compare text mining tools with UIMA
abstract
SUMMARY: Due to the increasing number of text mining resources (tools and corpora) available to biologists, interoperability issues between these resources are becoming significant obstacles to using them effectively. UIMA, the Unstructured Information Management Architecture, is an open framework designed to aid in the construction of more interoperable tools. U-Compare is built on top of the UIMA framework, and provides both a concrete framework for out-of-the-box text mining and a sophisticated evaluation platform allowing users to run specific tools on any target text, generating both detailed statistics and instance-based visualizations of outputs. U-Compare is a joint project, providing the world's largest, and still growing, collection of UIMA-compatible resources. These resources, originally developed by different groups for a variety of domains, include many famous tools and corpora. U-Compare can be launched straight from the web, without needing to be manually installed. All U-Compare components are provided ready-to-use and can be combined easily via a drag-and-drop interface without any programming. External UIMA components can also simply be mixed with U-Compare components, without distinguishing between locally and remotely deployed resources. AVAILABILITY: http://u-compare.org/
Yoshinobu Kano, William A. Baumgartner Jr., Luke McCrohon, Sophia Ananiadou, Kevin Cohen 0001, Lawrence Hunter, Jun'ichi Tsujii
Bioinform.6
2009 Ontology quality assurance through analysis of term transformations
abstract
MOTIVATION: It is important for the quality of biological ontologies that similar concepts be expressed consistently, or univocally. Univocality is relevant for the usability of the ontology for humans, as well as for computational tools that rely on regularity in the structure of terms. However, in practice terms are not always expressed consistently, and we must develop methods for identifying terms that are not univocal so that they can be corrected. RESULTS: We developed an automated transformation-based clustering methodology for detecting terms that use different linguistic conventions for expressing similar semantics. These term sets represent occurrences of univocality violations. Our method was able to identify 67 examples of univocality violations in the Gene Ontology. AVAILABILITY: The identified univocality violations are available upon request. We are preparing a release of an open source version of the software to be available at http://bionlp.sourceforge.net.
Karin Verspoor, Daniel Dvorkin, Kevin Cohen 0001, Lawrence Hunter
Bioinform.4
2009 Leveraging existing biological knowledge in the identification of candidate genes for facial dysmorphology
abstract
BACKGROUND: In response to the frequently overwhelming output of high-throughput microarray experiments, we propose a methodology to facilitate interpretation of biological data in the context of existing knowledge. Through the probabilistic integration of explicit and implicit data sources a functional interaction network can be constructed. Each edge connecting two proteins is weighted by a confidence value capturing the strength and reliability of support for that interaction given the combined data sources. The resulting network is examined in conjunction with expression data to identify groups of genes with significant temporal or tissue specific patterns. In contrast to unstructured gene lists, these networks often represent coherent functional groupings. RESULTS: By linking from shared functional categorizations to primary biological resources we apply this method to craniofacial microarray data, generating biologically testable hypotheses and identifying candidate genes for craniofacial development. CONCLUSION: The novel methodology presented here illustrates how the effective integration of pre-existing biological knowledge and high-throughput experimental data drives biological discovery and hypothesis generation.
Hannah J. Tipney, Sonia M. Leach, Weiguo Feng, Richard A. Spritz, Trevor Williams, Lawrence Hunter
BMC Bioinform.6
2009 The textual characteristics of traditional and Open Access scientific journals are similar
abstract
BACKGROUND: Recent years have seen an increased amount of natural language processing (NLP) work on full text biomedical journal publications. Much of this work is done with Open Access journal articles. Such work assumes that Open Access articles are representative of biomedical publications in general and that methods developed for analysis of Open Access full text publications will generalize to the biomedical literature as a whole. If this assumption is wrong, the cost to the community will be large, including not just wasted resources, but also flawed science. This paper examines that assumption. RESULTS: We collected two sets of documents, one consisting only of Open Access publications and the other consisting only of traditional journal publications. We examined them for differences in surface linguistic structures that have obvious consequences for the ease or difficulty of natural language processing and for differences in semantic content as reflected in lexical items. Regarding surface linguistic structures, we examined the incidence of conjunctions, negation, passives, and pronominal anaphora, and found that the two collections did not differ. We also examined the distribution of sentence lengths and found that both collections were characterized by the same mode. Regarding lexical items, we found that the Kullback-Leibler divergence between the two collections was low, and was lower than the divergence between either collection and a reference corpus. Where small differences did exist, log likelihood analysis showed that they were primarily in the area of formatting and in specific named entities. CONCLUSION: We did not find structural or semantic differences between the Open Access and traditional journal collections.
Karin Verspoor, Kevin Cohen 0001, Lawrence Hunter
BMC Bioinform.3
2009 New JBI emphasis on translational bioinformatics
Edward H. Shortliffe, Andrea Califano, Lawrence Hunter
J. Biomed. Informatics3
2009 Biomedical Discovery Acceleration, with Applications to Craniofacial Development
abstract
The profusion of high-throughput instruments and the explosion of new results in the scientific literature, particularly in molecular biomedicine, is both a blessing and a curse to the bench researcher. Even knowledgeable and experienced scientists can benefit from computational tools that help navigate this vast and rapidly evolving terrain. In this paper, we describe a novel computational approach to this challenge, a knowledge-based system that combines reading, reasoning, and reporting methods to facilitate analysis of experimental data. Reading methods extract information from external resources, either by parsing structured data or using biomedical language processing to extract information from unstructured data, and track knowledge provenance. Reasoning methods enrich the knowledge that results from reading by, for example, noting two genes that are annotated to the same ontology term or database entry. Reasoning is also used to combine all sources into a knowledge network that represents the integration of all sorts of relationships between a pair of genes, and to calculate a combined reliability score. Reporting methods combine the knowledge network with a congruent network constructed from experimental data and visualize the combined network in a tool that facilitates the knowledge-based analysis of that data. An implementation of this approach, called the Hanalyzer, is demonstrated on a large-scale gene expression array dataset relevant to craniofacial development. The use of the tool was critical in the creation of hypotheses regarding the roles of four genes never previously characterized as involved in craniofacial development; each of these hypotheses was validated by further experimental work.
Sonia M. Leach, Hannah J. Tipney, Weiguo Feng, William A. Baumgartner Jr., Priyanka Kasliwal, Ronald P. Schuyler, Trevor Williams, Richard A. Spritz, Lawrence Hunter
PLoS Comput. Biol.9
2008 Identification of OBO nonalignments and its implications for OBO enrichment
abstract
MOTIVATION: Existing projects that focus on the semiautomatic addition of links between existing terms in the Open Biomedical Ontologies can take advantage of reasoners that can make new inferences between terms that are based on the added formal definitions and that reflect nonalignments between the linked terms. However, these projects require that these definitions be necessary and sufficient, a strong requirement that often does not hold. If such definitions cannot be added, the reasoners cannot point to the nonalignments through the suggestion of new inferences. RESULTS: We describe a methodology by which we have identified over 1900 instances of nonredundant nonalignments between terms from the Gene Ontology (GO) biological process (BP), cellular component (CC) and molecular function (MF) ontologies, Chemical Entities of Biological Interest (ChEBI) and the Cell Type Ontology (CL). Many of the 39.8% of these nonalignments whose object terms are more atomic than the subject terms are not currently examined in other ontology-enrichment projects due to the fact that the necessary and sufficient conditions required for the inferences are not currently examined. Analysis of the ratios of nonalignments to assertions from which the nonalignments were identified suggests that BP-MF, BP-BP, BP-CL and CC-CC terms are relatively well-aligned, while ChEBI-MF, BP-ChEBI and CC-MF terms are relatively not aligned well. We propose four ways to resolve an identified nonalignment and recommend an analogous implementation of our methodology in ontology-enrichment tools to identify types of nonalignments that are currently not detected. AVAILABILITY: The nonalignments discussed in this article may be viewed at http://compbio.uchsc.edu/Hunter_lab/Bada/nonalignments_2008_03_06.html. Code for the generation of these nonalignments is available upon request. CONTACT: [email protected].
Michael Bada, Lawrence Hunter
Bioinform.2
2008 Semantic role labeling for protein transport predicates
abstract
BACKGROUND: Automatic semantic role labeling (SRL) is a natural language processing (NLP) technique that maps sentences to semantic representations. This technique has been widely studied in the recent years, but mostly with data in newswire domains. Here, we report on a SRL model for identifying the semantic roles of biomedical predicates describing protein transport in GeneRIFs - manually curated sentences focusing on gene functions. To avoid the computational cost of syntactic parsing, and because the boundaries of our protein transport roles often did not match up with syntactic phrase boundaries, we approached this problem with a word-chunking paradigm and trained support vector machine classifiers to classify words as being at the beginning, inside or outside of a protein transport role. RESULTS: We collected a set of 837 GeneRIFs describing movements of proteins between cellular components, whose predicates were annotated for the semantic roles AGENT, PATIENT, ORIGIN and DESTINATION. We trained these models with the features of previous word-chunking models, features adapted from phrase-chunking models, and features derived from an analysis of our data. Our models were able to label protein transport semantic roles with 87.6% precision and 79.0% recall when using manually annotated protein boundaries, and 87.0% precision and 74.5% recall when using automatically identified ones. CONCLUSION: We successfully adapted the word-chunking classification paradigm to semantic role labeling, applying it to a new domain with predicates completely absent from any previous studies. By combining the traditional word and phrasal role labeling features with biomedical features like protein boundaries and MEDPOST part of speech tags, we were able to address the challenges posed by the new domain data and subsequently build robust models that achieved F-measures as high as 83.1. This system for extracting protein transport information from GeneRIFs performs well even with proteins identified automatically, and is therefore more robust than the rule-based methods previously used to extract protein transport roles.
Steven Bethard, Zhiyong Lu, James H. Martin, Lawrence Hunter
BMC Bioinform.4
2008 Improving protein function prediction methods with integrated literature data
abstract
BACKGROUND: Determining the function of uncharacterized proteins is a major challenge in the post-genomic era due to the problem's complexity and scale. Identifying a protein's function contributes to an understanding of its role in the involved pathways, its suitability as a drug target, and its potential for protein modifications. Several graph-theoretic approaches predict unidentified functions of proteins by using the functional annotations of better-characterized proteins in protein-protein interaction networks. We systematically consider the use of literature co-occurrence data, introduce a new method for quantifying the reliability of co-occurrence and test how performance differs across species. We also quantify changes in performance as the prediction algorithms annotate with increased specificity. RESULTS: We find that including information on the co-occurrence of proteins within an abstract greatly boosts performance in the Functional Flow graph-theoretic function prediction algorithm in yeast, fly and worm. This increase in performance is not simply due to the presence of additional edges since supplementing protein-protein interactions with co-occurrence data outperforms supplementing with a comparably-sized genetic interaction dataset. Through the combination of protein-protein interactions and co-occurrence data, the neighborhood around unknown proteins is quickly connected to well-characterized nodes which global prediction algorithms can exploit. Our method for quantifying co-occurrence reliability shows superior performance to the other methods, particularly at threshold values around 10% which yield the best trade off between coverage and accuracy. In contrast, the traditional way of asserting co-occurrence when at least one abstract mentions both proteins proves to be the worst method for generating co-occurrence data, introducing too many false positives. Annotating the functions with greater specificity is harder, but co-occurrence data still proves beneficial. CONCLUSION: Co-occurrence data is a valuable supplemental source for graph-theoretic function prediction algorithms. A rapidly growing literature corpus ensures that co-occurrence data is a readily-available resource for nearly every studied organism, particularly those with small protein interaction databases. Though arguably biased toward known genes, co-occurrence data provides critical additional links to well-studied regions in the interaction network that graph-theoretic function prediction algorithms can exploit.
Aaron Gabow, Sonia M. Leach, William A. Baumgartner Jr., Lawrence Hunter, Debra Goldberg
BMC Bioinform.4
2008 OpenDMAP: An open source, ontology-driven concept analysis engine, with applications to capturing knowledge regarding protein transport, protein interactions and cell-type-specific gene expression
abstract
BACKGROUND: Information extraction (IE) efforts are widely acknowledged to be important in harnessing the rapid advance of biomedical knowledge, particularly in areas where important factual information is published in a diverse literature. Here we report on the design, implementation and several evaluations of OpenDMAP, an ontology-driven, integrated concept analysis system. It significantly advances the state of the art in information extraction by leveraging knowledge in ontological resources, integrating diverse text processing applications, and using an expanded pattern language that allows the mixing of syntactic and semantic elements and variable ordering. RESULTS: OpenDMAP information extraction systems were produced for extracting protein transport assertions (transport), protein-protein interaction assertions (interaction) and assertions that a gene is expressed in a cell type (expression). Evaluations were performed on each system, resulting in F-scores ranging from .26-.72 (precision .39-.85, recall .16-.85). Additionally, each of these systems was run over all abstracts in MEDLINE, producing a total of 72,460 transport instances, 265,795 interaction instances and 176,153 expression instances. CONCLUSION: OpenDMAP advances the performance standards for extracting protein-protein interaction predications from the full texts of biomedical research articles. Furthermore, this level of performance appears to generalize to other information extraction tasks, including extracting information about predicates of more than two arguments. The output of the information extraction system is always constructed from elements of an ontology, ensuring that the knowledge representation is grounded with respect to a carefully constructed model of reality. The results of these efforts can be used to increase the efficiency of manual curation efforts and to provide additional features in systems that integrate multiple sources for information extraction. The open source OpenDMAP code library is freely available at http://bionlp.sourceforge.net/
Lawrence Hunter, Zhiyong Lu, James Firby, William A. Baumgartner Jr., Helen L. Johnson 0001, Philip V. Ogren, Kevin Cohen 0001
BMC Bioinform.1
2008 Predicting protein linkages in bacteria: Which method is best depends on task
abstract
BACKGROUND: Applications of computational methods for predicting protein functional linkages are increasing. In recent years, several bacteria-specific methods for predicting linkages have been developed. The four major genomic context methods are: Gene cluster, Gene neighbor, Rosetta Stone, and Phylogenetic profiles. These methods have been shown to be powerful tools and this paper provides guidelines for when each method is appropriate by exploring different features of each method and potential improvements offered by their combination. We also review many previous treatments of these prediction methods, use the latest available annotations, and offer a number of new observations. RESULTS: Using Escherichia coli K12 and Bacillus subtilis, linkage predictions made by each of these methods were evaluated against three benchmarks: functional categories defined by COG and KEGG, known pathways listed in EcoCyc, and known operons listed in RegulonDB. Each evaluated method had strengths and weaknesses, with no one method dominating all aspects of predictive ability studied. For functional categories, as previous studies have shown, the Rosetta Stone method was individually best at detecting linkages and predicting functions among proteins with shared KEGG categories while the Phylogenetic profile method was best for linkage detection and function prediction among proteins with common COG functions. Differences in performance under COG versus KEGG may be attributable to the presence of paralogs. Better function prediction was observed when using a weighted combination of linkages based on reliability versus using a simple unweighted union of the linkage sets. For pathway reconstruction, 99 complete metabolic pathways in E. coli K12 (out of the 209 known, non-trivial pathways) and 193 pathways with 50% of their proteins were covered by linkages from at least one method. Gene neighbor was most effective individually on pathway reconstruction, with 48 complete pathways reconstructed. For operon prediction, Gene cluster predicted completely 59% of the known operons in E. coli K12 and 88% (333/418)in B. subtilis. Comparing two versions of the E. coli K12 operon database, many of the unannotated predictions in the earlier version were updated to true predictions in the later version. Using only linkages found by both Gene Cluster and Gene Neighbor improved the precision of operon predictions. Additionally, as previous studies have shown, combining features based on intergenic region and protein function improved the specificity of operon prediction. CONCLUSION: A common problem for computational methods is the generation of a large number of false positives that might be caused by an incomplete source of validation. By comparing two versions of a database, we demonstrated the dramatic differences on reported results. We used several benchmarks on which we have shown the comparative effectiveness of each prediction method, as well as provided guidelines as to which method is most appropriate for a given prediction task.
Anis Karimpour-Fard, Sonia M. Leach, Ryan T. Gill, Lawrence Hunter
BMC Bioinform.4
2008 Getting Started in Text Mining
abstract
Text mining is the use of automated methods for exploiting the enormous amount of knowledge available in the biomedical literature. There are at least as many motivations for doing text mining work as there are types of bioscientists. Model organism database curators have been heavy participants in the development of the field due to their need to process large numbers of publications in order to populate the many data fields for every gene in their species of interest. Bench scientists have built biomedical text mining applications to aid in the development of tools for interpreting the output of high-throughput assays and to improve searches of sequence databases (see [1] for a review). Bioscientists of every stripe have built applications to deal with the dual issues of the double-exponential growth in the scientific literature over the past few years and of the unique issues in searching PubMed/MEDLINE for genomics-related publications. A surprising phenomenon can be noted in the recent history of biomedical text mining: although several systems have been built and deployed in the past few years—Chilibot, Textpresso, and PreBIND (see Text S1 for these and most other citations), for example—the ones that are seeing high usage rates and are making productive contributions to the working lives of bioscientists have been built not by text mining specialists, but by bioscientists. We speculate on why this might be so below. Three basic types of approaches to text mining have been prevalent in the biomedical domain. Co-occurrence–based methods do no more than look for concepts that occur in the same unit of text—typically a sentence, but sometimes as large as an abstract—and posit a relationship between them. (See [2] for an early co-occurrence–based system.) For example, if such a system saw that BRCA1 and breast cancer occurred in the same sentence, it might assume a relationship between breast cancer and the BRCA1 gene. Some early biomedical text mining systems were co-occurrence–based, but such systems are highly error prone, and are not commonly built today. In fact, many text mining practitioners would not consider them to be text mining systems at all. Co-occurrence of concepts in a text is sometimes used as a simple baseline when evaluating more sophisticated systems; as such, they are nontrivial, since even a co-occurrence–based system must deal with variability in the ways that concepts are expressed in human-produced texts. For example, BRCA1 could be referred to by any of its alternate symbols—IRIS, PSCP, BRCAI, BRCC1, or RNF53 (or by any of their many spelling variants, which include BRCA1, BRCA-1, and BRCA 1)—or by any of the variants of its full name, viz. breast cancer 1, early onset (its official name per Entrez Gene and the Human Gene Nomenclature Committee), as breast cancer susceptibility gene 1, or as the latter's variant breast cancer susceptibility gene-1. Similarly, breast cancer could be referred to as breast cancer, carcinoma of the breast, or mammary neoplasm. These variability issues challenge more sophisticated systems, as well; we discuss ways of coping with them in Text S1. Two more common (and more sophisticated) approaches to text mining exist: rule-based or knowledge-based approaches, and statistical or machine-learning-based approaches. The variety of types of rule-based systems is quite wide. In general, rule-based systems make use of some sort of knowledge. This might take the form of general knowledge about how language is structured, specific knowledge about how biologically relevant facts are stated in the biomedical literature, knowledge about the sets of things that bioscientists talk about and the kinds of relationships that they can have with one another, and the variant forms by which they might be mentioned in the literature, or any subset or combination of these. (See [3] for an early rule-based system, and [4] for a discussion of rule-based approaches to various biomedical text mining tasks.) At one end of the spectrum, a simple rule-based system might use hard-coded patterns—for example, plays a role in or is associated with —to find explicit statements about the classes of things in which the researcher is interested. At the other end of the spectrum, a rule-based system might use sophisticated linguistic and semantic analyses to recognize a wide range of possible ways of making assertions about those classes of things. It is worth noting that useful systems have been built using technologies at both ends of the spectrum, and at many points in between. In contrast, statistical or machine-learning–based systems operate by building classifiers that may operate on any level, from labelling part of speech to choosing syntactic parse trees to classifying full sentences or documents. (See [5] for an early learning-based system, and [4] for a discussion of learning-based approaches to various biomedical text mining tasks.) Rule-based and statistical systems each have their advantages and disadvantages. For example, rule systems are often assumed (not necessarily correctly) to take a significant amount of time to develop. Statistical systems typically require large amounts of expensive-to-get labelled training data. In practice, statistical and rule-based systems can be fruitfully combined. For example, a statistical system that classifies documents as to whether or not they are relevant to the subject of genetic variation in mouse genes might use the output of a rule-based mutation recognizer as one of its feature extractors. Many systems also employ an initial statistical processing step, followed by rule-based post-processing. A primary problem that either type of system must deal with is the issue of ambiguity: the existence of multiple relationships between language and meanings or categories. Ambiguity exists at every level of linguistic structure, from the part of speech of words to subtle issues in pragmatics. A common example of ambiguity in genomics text is related to gene names and symbols. Consider the string fat: is it an adjective, or a noun? Either part of speech is entirely plausible in biomedical texts, and PubMed returns almost 112 K hits for that single-word query (and more than 13 K even if we try to restrict the query to genomics by including the disjunction (gene OR genetic OR genetics). This ambiguity is relatively easy to resolve, but fat also turns out to be the name or symbol of a number of different genes—humans, mice, rats, Drosophila, zebrafish, chickens, M. mulatta, and two Lactobacilli have at least one gene whose name, official symbol, or alias is fat. Even if the species whose gene is being referred to can be determined, the ambiguity may still not be resolved—in humans, fat is the official symbol of Entrez Gene entry 2195 and an alternate symbol for Entrez Gene entry 948. The distinction is not trivial. The former is a cadhedrin, and is associated with tumor suppression and with bipolar disorder, while the latter is a thrombospondin receptor associated with atherosclerosis, platelet glycoprotein deficiency, hyperlipidemia, and insulin resistance, to name just a few phenotypes. These ambiguities are not trivial: if your analysis is wrong, you miss or erroneously extract information on relations between molecular biology and human disease.
Kevin Cohen 0001, Lawrence Hunter
PLoS Comput. Biol.2
2007 Mining Discriminative Distance Context of Transcription Factor Binding Sites on ChIP Enriched Regions
Katerina J. Kechris, Lawrence Hunter
ISBRA3
2007 MutationFinder: a high-performance system for extracting point mutation mentions from text
abstract
Discussion of point mutations is ubiquitous in biomedical literature, and manually compiling databases or literature on mutations in specific genes or proteins is tedious. We present an open-source, rule-based system, MutationFinder, for extracting point mutation mentions from text. On blind test data, it achieves nearly perfect precision and a markedly improved recall over a baseline. AVAILABILITY: MutationFinder, along with a high-quality gold standard data set, and a scoring script for mutation extraction systems have been made publicly available. Implementations, source code and unit tests are available in Python, Perl and Java. MutationFinder can be used as a stand-alone script, or imported by other applications. PROJECT URL: http://bionlp.sourceforge.net.
J. Gregory Caporaso, William A. Baumgartner Jr., David A. Randolph, Kevin Cohen 0001, Lawrence Hunter
Bioinform.5
2007 Enrichment of OBO ontologies
Michael Bada, Lawrence Hunter
J. Biomed. Informatics2
2007 The International Society for Computational Biology 10th Anniversary
abstract
PLoS Computational Biology is the official journal of the International Society for Computational Biology (ISCB), a partnership that was formed during the Journal's conception in 2005. With ISCB being the only international body representing computational biologists, it made perfect sense for PLoS Computational Biology to be closely affiliated. The Society had to take more of a chance than similar societies, choosing to step away from an existing financially beneficial subscription journal to align with an open access publication as a matter of principle. To our knowledge, ISCB was the first major international scientific society to do so. Now, as PLoS Computational Biology reaches its two-year mark, ISCB simultaneously celebrates its tenth anniversary, having formed officially on June 18, 1997. We early presidents of ISCB reflect on the state of computational biology ten years ago, how far we have come since, and what thought-provoking future challenges might lie ahead with regard to innovations in publishing technologies.
Lawrence Hunter, Russ B. Altman, Philip E. Bourne
PLoS Comput. Biol.1
2006 A critical review of PASBio's argument structures for biomedical verbs
abstract
BACKGROUND: Propositional representations of biomedical knowledge are a critical component of most aspects of semantic mining in biomedicine. However, the proper set of propositions has yet to be determined. Recently, the PASBio project proposed a set of propositions and argument structures for biomedical verbs. This initial set of representations presents an opportunity for evaluating the suitability of predicate-argument structures as a scheme for representing verbal semantics in the biomedical domain. Here, we quantitatively evaluate several dimensions of the initial PASBio propositional structure repository. RESULTS: We propose a number of metrics and heuristics related to arity, role labelling, argument realization, and corpus coverage for evaluating large-scale predicate-argument structure proposals. We evaluate the metrics and heuristics by applying them to PASBio 1.0. CONCLUSION: PASBio demonstrates the suitability of predicate-argument structures for representing aspects of the semantics of biomedical verbs. Metrics related to theta-criterion violations and to the distribution of arguments are able to detect flaws in semantic representations, given a set of predicate-argument structures and a relatively small corpus annotated with them.
Kevin Cohen 0001, Lawrence Hunter
BMC Bioinform.2
2005 Empirical data on corpus design and usage in biomedical natural language processing
Kevin Cohen 0001, Philip V. Ogren, Lynne M. Fox, Lawrence Hunter
AMIA4
2005 BioCreAtIvE Task1A: entity identification with a stochastic tagger
abstract
BACKGROUND: Our approach to Task 1A was inspired by Tanabe and Wilbur's ABGene system. Like Tanabe and Wilbur, we approached the problem as one of part-of-speech tagging, adding a GENE tag to the standard tag set. Where their system uses the Brill tagger, we used TnT, the Trigrams 'n' Tags HMM-based part-of-speech tagger. Based on careful error analysis, we implemented a set of post-processing rules to correct both false positives and false negatives. We participated in both the open and the closed divisions; for the open division, we made use of data from NCBI. RESULTS: Our base system without post-processing achieved a precision and recall of 68.0% and 77.2%, respectively, giving an F-measure of 72.3%. The full system with post-processing achieved a precision and recall of 80.3% and 80.5% giving an F-measure of 80.4%. We achieved a slight improvement (F-measure = 80.9%) by employing a dictionary-based post-processing step for the open division. We placed third in both the open and the closed division. CONCLUSION: Our results show that a part-of-speech tagger can be augmented with post-processing rules resulting in an entity identification system that competes well with other approaches.
Shuhei Kinoshita, Kevin Cohen 0001, Philip V. Ogren, Lawrence Hunter
BMC Bioinform.4
2002 A bioinformatics tool to select sequences for microarray studies of mouse models of oncogenesis
abstract
Abstract One of the challenges to the effective utilization of cDNA microarray analysis in mouse models of oncogenesis is the choice of a critical set of probes that are informative for human disease. Given the thousands of genes with a potential role in human oncogenesis and the hundreds of thousands of mouse sequences available for use as probes, selection of an informative set of mouse probes can be an overwhelming task. We have developed a web based sequence mining tool using DataBase Independent (DBI) Perl to annotate publicly available sequences. The Mouse Oncochip Design Tool uses the Mouse Genome Database (MGD) developed and maintained by the Jackson Laboratories for mouse DNA sequences. There are over 380 000 sequences in their database. The output list has been ordered to present the genes more likely to be informative in a mouse model of human cancer using a candidate set of oncogenes to order the list. Mouse sequences that represent genes that are homologous with a member of a human oncogene set are listed first. In addition it provides a set of links for information on clone source gene function. Contact: http://nciarray.nci.nih.gov/cgi-bin/me/mouse_design.cgi * To whom correspondence should be addressed. 5 Current address: Department of Pathology and Biomedical Informatics, Vanderbilt University Medical Center, MCN C-3321, Nashville TN 37232, USA. 6 Current address: Department of Pharmacology, School of Medicine, University of Colorado Health Sciences Center, 4200 E Ninth Avenue, Denver CO 80262, USA. 7 Current address: Genome Institute of Singapore, 1 Research Link, IMA Building #04-01, National University of Singapore, Singapore 117604.
Mary E. Edgerton, Ronald C. Taylor, John I. Powell, Lawrence Hunter, Richard Simon, Edison T. Liu
Bioinform.4
1999 Mining molecular binding terminology from biomedical text
Thomas C. Rindflesch, Lawrence Hunter, Alan R. Aronson
AMIA2
1998 Identification of Divergent Functions in Homologous Proteins by Induction over Conserved Modules
Imran Shah, Lawrence Hunter
ISMB2
1997 Identifying Chimerism in Proteins Using Hidden Markov Models of Codon Usage
Lawrence Hunter, Barry Zeeberg
ISMB1
1997 Predicting Enzyme Function from Sequence: A Systematic Appraisal
Imran Shah, Lawrence Hunter
ISMB2
1995 Introduction
Jude W. Shavlik, Lawrence Hunter, David B. Searls
Mach. Learn.2
1993 Grand Challenge AI Applications
Hiroaki Kitano, Walther von Hahn, Lawrence Hunter, Benjamin W. Wah, Toshio Yokoi
IJCAI3
1993 Finding Relevant Biomolecular Features
Lawrence Hunter, Teri E. Klein
ISMB1
1993 Computationally Efficient Cluster Representation in Molecular Sequence Megaclassification
David J. States, Nomi L. Harris, Lawrence Hunter
ISMB3
1992 Mega-Classification: Discovering Motifs in Massive Datastreams
Nomi L. Harris, Lawrence Hunter, David J. States
AAAI2
1992 Artificial Intelligence and Molecular Biology
Lawrence Hunter
AAAI1
1992 Efficient Classification of Massive, Unsegmented Datastreams
Lawrence Hunter, Nomi L. Harris, David J. States
ML1
1992 The use of explicit goals for knowledge to guide inference and learning
Ashwin Ram 0001, Lawrence Hunter
Appl. Intell.2
1991 A Goal-Based Approach to Intelligent Information Retrieval
Ashwin Ram 0001, Lawrence Hunter
ML2
1989 Knowledge Acquisition Planning: Results and Prospects
Lawrence Hunter
ML1