VLDB 2026 Research / reviewers in the wild / expert
Russ B. Altman
dblp:47/2081
· DBLP profile ↗
136ranked-venue papers
23as first author
16since 2021 · last 2025
0000-0003-3859-2905ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 124 · 18 first-author · 15 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 1 since 2021Systems, architecture and hardware · 3Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Unsupervised learning reveals landscape of local structural motifs across protein classesabstractMOTIVATION: Proteins are known to share similarities in local regions of three-dimensional (3D) structure even across disparate global folds. Such correspondences can help to shed light on functional relationships between proteins and identify conserved local structural features that lead to function. Self-supervised deep learning on large protein structure datasets has produced high-fidelity representations of local structural microenvironments, providing the opportunity to characterize the landscape of local structure and function at scale. RESULTS: In this work, we leverage these representations to cluster over 15 million environments in the Protein Data Bank, resulting in the creation of a "lexicon" of local 3D motifs which form the building blocks of all known protein structures. We characterize these motifs and demonstrate that they provide valuable information for modeling structure and function at all scales of protein analysis, from full protein chains to binding pockets to individual amino acids. We devise a new protein representation based solely on its constituent local motifs and show that this representation enables state-of-the-art performance on protein structure search and model quality assessment. We then show that this approach enables accurate prediction of drug off-target interactions by modeling the similarity between local binding pockets. Finally, we identify structural motifs associated with pathogenic variants in the human proteome by leveraging the predicted structures in the AlphaFold structure database. AVAILABILITY AND IMPLEMENTATION: All code and cluster data are available at https://github.com/awfderry/collapse-motifs. Alexander Derry, Haim Krupkin, Alp Tartici, Russ B. Altman |
Bioinform. | 4 |
| 2025 | Semi-supervised data-integrated feature importance enhances performance and interpretability of biological classification tasksabstractMOTIVATION: Accurate model performance on training data does not ensure alignment between the model's feature weighting patterns and human knowledge, which can limit the model's relevance and applicability. We propose Semi-Supervised Data-Integrated Feature Importance (DIFI), a method that numerically integrates a priori knowledge, represented as a sparse knowledge map, into the model's feature weighting. By incorporating the similarity between the knowledge map and the feature map into a loss function, DIFI causes the model's feature weighting to correlate with the knowledge. RESULTS: We show that DIFI can improve the performance of neural networks using two biological tasks. In the first task, cancer type prediction from gene expression profiles was guided by identities of cancer type-specific biomarkers. In the second task, enzyme/non-enzyme classification from protein sequences was guided by the locations of the catalytic residues. In both tasks, DIFI leads to improved performance and feature weighting that is interpretable. DIFI is a novel method for injecting knowledge to achieve model alignment and interpretability. AVAILABILITY AND IMPLEMENTATION: Code and models for our experiments are available at https://github.com/junwkim1/DIFI. Jun W. Kim, Russ B. Altman |
Bioinform. | 2 |
| 2025 | Pool PaRTI: a PageRank-based pooling method for identifying critical residues and enhancing protein sequence representationsabstractMOTIVATION: Protein language models (PLMs) produce token-level embeddings for each residue, resulting in an output matrix with dimensions that vary based on sequence length. However, downstream machine learning models typically require fixed-length input vectors, necessitating a pooling method to compress the output matrix into a single vector representation of the entire protein. Traditional pooling methods often result in substantial information loss, impacting downstream task performance. We aim to develop a pooling method that produces more expressive general-purpose protein embedding vectors while offering biological interpretability. RESULTS: We introduce Pool PaRTI, a novel pooling method that leverages internal transformer attention matrices and PageRank to assign token importance weights. Our unsupervised and parameter-free approach consistently prioritizes residues experimentally annotated as critical for function, assigning them higher importance scores. Across four diverse protein machine learning tasks, Pool PaRTI enables significant performance gains in predictive performance. Additionally, it enhances interpretability by identifying biologically relevant regions without relying on explicit structural data or annotated training. To assess generalizability, we evaluated Pool PaRTI with two encoder-only PLMs, confirming its robustness across different models. AVAILABILITY AND IMPLEMENTATION: Pool PaRTI is implemented in Python with PyTorch and is available at github.com/Helix-Research-Lab/Pool_PaRTI.git. The Pool PaRTI sequence embeddings and residue importance values for all human proteins on UniProt are available at zenodo.org/records/15036725 for ESM2 and protBERT. Alp Tartici, Gowri Nayar, Russ B. Altman |
Bioinform. | 3 |
| 2025 | Paying attention to attention: High attention sites as indicators of protein family and function in language modelsabstractProtein Language Models (PLMs) use transformer architectures to capture patterns within protein primary sequences, providing a powerful computational representation of the amino acid sequence. Through large-scale training on protein primary sequences, PLMs generate vector representations that encapsulate the biochemical and structural properties of proteins. At the core of PLMs is the attention mechanism, which facilitates the capture of long-range dependencies by computing pairwise importance scores across residues, thereby highlighting regions of biological interaction within the sequence. The attention matrices offer an untapped opportunity to uncover specific biological properties of proteins, particularly their functions. In this work, we introduce a novel approach, using the Evolutionary Scale Modelling (ESM), for identifying High Attention (HA) sites within protein primary sequences, corresponding to key residues that define protein families. By examining attention patterns across multiple layers, we pinpoint residues that contribute most to family classification and function prediction. Our contributions are as follows: (1) we propose a method for identifying HA sites at critical residues from the middle layers of the PLM; (2) we demonstrate that these HA sites provide interpretable links to biological functions; and (3) we show that HA sites improve active site predictions for functions of unannotated proteins. We make available the HA sites for the human proteome. This work offers a broadly applicable approach to protein classification and functional annotation and provides a biological interpretation of the PLM's representation. Gowri Nayar, Alp Tartici, Russ B. Altman |
PLoS Comput. Biol. | 3 |
| 2024 | Prospector Heads: Generalized Feature Attribution for Large Models & DataabstractFeature attribution, the ability to localize regions of the input data that are relevant for classification, is an important capability for ML models in scientific and biomedical domains. Current methods for feature attribution, which rely on "explaining" the predictions of end-to-end classifiers, suffer from imprecise feature localization and are inadequate for use with small sample sizes and high-dimensional datasets due to computational challenges. We introduce prospector heads, an efficient and interpretable alternative to explanation-based attribution methods that can be applied to any encoder and any data modality. Prospector heads generalize across modalities through experiments on sequences (text), images (pathology), and graphs (protein structures), outperforming baseline attribution methods by up to 26.3 points in mean localization AUPRC. We also demonstrate how prospector heads enable improved interpretation and discovery of class-specific patterns in input data. Through their high performance, flexibility, and generalizability, prospectors provide a framework for improving trust and transparency for ML models in complex domains. Gautam Machiraju, Alexander Derry, Arjun D. Desai, Neel Guha, James Zou 0001, Russ B. Altman, Christopher Ré, Parag Mallick |
ICML | 7 |
| 2023 | Gene set proximity analysis: expanding gene set enrichment analysis through learned geometric embeddings, with drug-repurposing applications in COVID-19abstractMOTIVATION: Gene set analysis methods rely on knowledge-based representations of genetic interactions in the form of both gene set collections and protein-protein interaction (PPI) networks. However, explicit representations of genetic interactions often fail to capture complex interdependencies among genes, limiting the analytic power of such methods. RESULTS: We propose an extension of gene set enrichment analysis to a latent embedding space reflecting PPI network topology, called gene set proximity analysis (GSPA). Compared with existing methods, GSPA provides improved ability to identify disease-associated pathways in disease-matched gene expression datasets, while improving reproducibility of enrichment statistics for similar gene sets. GSPA is statistically straightforward, reducing to a version of traditional gene set enrichment analysis through a single user-defined parameter. We apply our method to identify novel drug associations with SARS-CoV-2 viral entry. Finally, we validate our drug association predictions through retrospective clinical analysis of claims data from 8 million patients, supporting a role for gabapentin as a risk factor and metformin as a protective factor for severe COVID-19. AVAILABILITY AND IMPLEMENTATION: GSPA is available for download as a command-line Python package at https://github.com/henrycousins/gspa. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Henry Cousins, Taryn Hall, Yinglong Guo, Luke Tso, Kathy T. H. Tzeng, Le Cong, Russ B. Altman |
Bioinform. | 7 |
| 2023 | POPDx: an automated framework for patient phenotyping across 392 246 individuals in the UK Biobank studyabstractOBJECTIVE: For the UK Biobank, standardized phenotype codes are associated with patients who have been hospitalized but are missing for many patients who have been treated exclusively in an outpatient setting. We describe a method for phenotype recognition that imputes phenotype codes for all UK Biobank participants. MATERIALS AND METHODS: POPDx (Population-based Objective Phenotyping by Deep Extrapolation) is a bilinear machine learning framework for simultaneously estimating the probabilities of 1538 phenotype codes. We extracted phenotypic and health-related information of 392 246 individuals from the UK Biobank for POPDx development and evaluation. A total of 12 803 ICD-10 diagnosis codes of the patients were converted to 1538 phecodes as gold standard labels. The POPDx framework was evaluated and compared to other available methods on automated multiphenotype recognition. RESULTS: POPDx can predict phenotypes that are rare or even unobserved in training. We demonstrate substantial improvement of automated multiphenotype recognition across 22 disease categories, and its application in identifying key epidemiological features associated with each phenotype. CONCLUSIONS: POPDx helps provide well-defined cohorts for downstream studies. It is a general-purpose method that can be applied to other biobanks with diverse but incomplete data. Sheng Wang 0012, Russ B. Altman |
J. Am. Medical Informatics Assoc. | 3 |
| 2023 | Associating biological context with protein-protein interactions through text mining at PubMed scale
Daniel N. Sosa, Rogier Hintzen, Betty Xiong, Alex de Giorgio, Julien Fauqueur, Mark Davies, Jake Lever, Russ B. Altman |
J. Biomed. Informatics | 8 |
| 2022 | Challenges and opportunities in network-based solutions for biological questionsabstractNetwork biology is useful for modeling complex biological phenomena; it has attracted attention with the advent of novel graph-based machine learning methods. However, biological applications of network methods often suffer from inadequate follow-up. In this perspective, we discuss obstacles for contemporary network approaches-particularly focusing on challenges representing biological concepts, applying machine learning methods, and interpreting and validating computational findings about biology-in an effort to catalyze actionable biological discovery. Margaret G. Guo, Daniel N. Sosa, Russ B. Altman |
Briefings Bioinform. | 3 |
| 2022 | Contexts and contradictions: a roadmap for computational drug repurposing with knowledge inferenceabstractThe cost of drug development continues to rise and may be prohibitive in cases of unmet clinical need, particularly for rare diseases. Artificial intelligence-based methods are promising in their potential to discover new treatment options. The task of drug repurposing hypothesis generation is well-posed as a link prediction problem in a knowledge graph (KG) of interacting of drugs, proteins, genes and disease phenotypes. KGs derived from biomedical literature are semantically rich and up-to-date representations of scientific knowledge. Inference methods on scientific KGs can be confounded by unspecified contexts and contradictions. Extracting context enables incorporation of relevant pharmacokinetic and pharmacodynamic detail, such as tissue specificity of interactions. Contradictions in biomedical KGs may arise when contexts are omitted or due to contradicting research claims. In this review, we describe challenges to creating literature-scale representations of pharmacological knowledge and survey current approaches toward incorporating context and resolving contradictions. Daniel N. Sosa, Russ B. Altman |
Briefings Bioinform. | 2 |
| 2022 | Construction of disease-specific cytokine profiles by associating disease genes with immune responsesabstractThe pathogenesis of many inflammatory diseases is a coordinated process involving metabolic dysfunctions and immune response-usually modulated by the production of cytokines and associated inflammatory molecules. In this work, we seek to understand how genes involved in pathogenesis which are often not associated with the immune system in an obvious way communicate with the immune system. We have embedded a network of human protein-protein interactions (PPI) from the STRING database with 14,707 human genes using feature learning that captures high confidence edges. We have found that our predicted Association Scores derived from the features extracted from STRING's high confidence edges are useful for predicting novel connections between genes, thus enabling the construction of a full map of predicted associations for all possible pairs between 14,707 human genes. In particular, we analyzed the pattern of associations for 126 cytokines and found that the six patterns of cytokine interaction with human genes are consistent with their functional classifications. To define the disease-specific roles of cytokines we have collected gene sets for 11,944 diseases from DisGeNET. We used these gene sets to predict disease-specific gene associations with cytokines by calculating the normalized average Association Scores between disease-associated gene sets and the 126 cytokines; this creates a unique profile of inflammatory genes (both known and predicted) for each disease. We validated our predicted cytokine associations by comparing them to known associations for 171 diseases. The predicted cytokine profiles correlate (p-value<0.0003) with the known ones in 95 diseases. We further characterized the profiles of each disease by calculating an "Inflammation Score" that summarizes different modes of immune responses. Finally, by analyzing subnetworks formed between disease-specific pathogenesis genes, hormones, receptors, and cytokines, we identified the key genes responsible for interactions between pathogenesis and inflammatory responses. These genes and the corresponding cytokines used by different immune disorders suggest unique targets for drug discovery. Tianyun Liu, Shiyin Wang, Michael Wornow, Russ B. Altman |
PLoS Comput. Biol. | 4 |
| 2021 | Randomized user testing of recommender system clinical decision support
Andre Kumar, Rachael C. Aikens, Jason Horn, Lisa Shieh, Mark A. Musen, Michael T. M. Baiocchi, Russ B. Altman, Mary K. Goldstein, Steven M. Asch, Jonathan H. Chen |
AMIA | 7 |
| 2021 | Large-scale labeling and assessment of sex bias in publicly available expression dataabstractBACKGROUND: Women are at more than 1.5-fold higher risk for clinically relevant adverse drug events. While this higher prevalence is partially due to gender-related effects, biological sex differences likely also impact drug response. Publicly available gene expression databases provide a unique opportunity for examining drug response at a cellular level. However, missingness and heterogeneity of metadata prevent large-scale identification of drug exposure studies and limit assessments of sex bias. To address this, we trained organism-specific models to infer sample sex from gene expression data, and used entity normalization to map metadata cell line and drug mentions to existing ontologies. Using this method, we inferred sex labels for 450,371 human and 245,107 mouse microarray and RNA-seq samples from refine.bio. RESULTS: Overall, we find slight female bias (52.1%) in human samples and (62.5%) male bias in mouse samples; this corresponds to a majority of mixed sex studies in humans and single sex studies in mice, split between female-only and male-only (25.8% vs. 18.9% in human and 21.6% vs. 31.1% in mouse, respectively). In drug studies, we find limited evidence for sex-sampling bias overall; however, specific categories of drugs, including human cancer and mouse nervous system drugs, are enriched in female-only and male-only studies, respectively. We leverage our expression-based sex labels to further examine the complexity of cell line sex and assess the frequency of metadata sex label misannotations (2-5%). CONCLUSIONS: Our results demonstrate limited overall sex bias, while highlighting high bias in specific subfields and underscoring the importance of including sex labels to better understand the underlying biology. We make our inferred and normalized labels, along with flags for misannotated samples, publicly available to catalyze the routine use of sex as a study variable in future analyses. Emily R. Flynn, Annie Chang, Russ B. Altman |
BMC Bioinform. | 3 |
| 2021 | Repurposing biomedical informaticians for COVID-19
Daniel N. Sosa, Binbin Chen 0002, Amit Kaushal, Adam Lavertu, Jake Lever, Stefano E. Rensi, Russ B. Altman |
J. Biomed. Informatics | 7 |
| 2021 | Search and visualization of gene-drug-disease interactions for pharmacogenomics and precision medicine research using GeneDiveabstractBACKGROUND: Understanding the relationships between genes, drugs, and disease states is at the core of pharmacogenomics. Two leading approaches for identifying these relationships in medical literature are: human expert led manual curation efforts, and modern data mining based automated approaches. The former generates small amounts of high-quality data, and the latter offers large volumes of mixed quality data. The algorithmically extracted relationships are often accompanied by supporting evidence, such as, confidence scores, source articles, and surrounding contexts (excerpts) from the articles, that can be used as data quality indicators. Tools that can leverage these quality indicators to help the user gain access to larger and high-quality data are needed. APPROACH: We introduce GeneDive, a web application for pharmacogenomics researchers and precision medicine practitioners that makes gene, disease, and drug interactions data easily accessible and usable. GeneDive is designed to meet three key objectives: (1) provide functionality to manage information-overload problem and facilitate easy assimilation of supporting evidence, (2) support longitudinal and exploratory research investigations, and (3) offer integration of user-provided interactions data without requiring data sharing. RESULTS: GeneDive offers multiple search modalities, visualizations, and other features that guide the user efficiently to the information of their interest. To facilitate exploratory research, GeneDive makes the supporting evidence and context for each interaction readily available and allows the data quality threshold to be controlled by the user as per their risk tolerance level. The interactive search-visualization loop enables relationship discoveries between diseases, genes, and drugs that might not be explicitly described in literature but are emergent from the source medical corpus and deductive reasoning. The ability to utilize user's data either in combination with the GeneDive native datasets or in isolation promotes richer data-driven exploration and discovery. These functionalities along with GeneDive's applicability for precision medicine, bringing the knowledge contained in biomedical literature to bear on particular clinical situations and improving patient care, are illustrated through detailed use cases. CONCLUSION: GeneDive is a comprehensive, broad-use biological interactions browser. The GeneDive application and information about its underlying system architecture are available at http://www.genedive.net. GeneDive Docker image is also available for download at this URL, allowing users to (1) import their own interaction data securely and privately; and (2) generate and test hypotheses across their own and other datasets. Mike Wong 0001, Paul Previde, Jack Cole, Brook Thomas, Nayana Laxmeshwar, Emily K. Mallory, Jake Lever, Dragutin Petkovic, Russ B. Altman, Anagha Kulkarni 0001 |
J. Biomed. Informatics | 9 |
| 2021 | Modeling drug response using network-based personalized treatment prediction (NetPTP) with applications to inflammatory bowel diseaseabstractFor many prevalent complex diseases, treatment regimens are frequently ineffective. For example, despite multiple available immunomodulators and immunosuppressants, inflammatory bowel disease (IBD) remains difficult to treat. Heterogeneity in the disease across patients makes it challenging to select the optimal treatment regimens, and some patients do not respond to any of the existing treatment choices. Drug repurposing strategies for IBD have had limited clinical success and have not typically offered individualized patient-level treatment recommendations. In this work, we present NetPTP, a Network-based Personalized Treatment Prediction framework which models measured drug effects from gene expression data and applies them to patient samples to generate personalized ranked treatment lists. To accomplish this, we combine publicly available network, drug target, and drug effect data to generate treatment rankings using patient data. These ranked lists can then be used to prioritize existing treatments and discover new therapies for individual patients. We demonstrate how NetPTP captures and models drug effects, and we apply our framework to individual IBD samples to provide novel insights into IBD treatment. Lichy Han, Zahra N. Sayyid, Russ B. Altman |
PLoS Comput. Biol. | 3 |
| 2020 | Extracting chemical reactions from text using SnorkelabstractBACKGROUND: Enzymatic and chemical reactions are key for understanding biological processes in cells. Curated databases of chemical reactions exist but these databases struggle to keep up with the exponential growth of the biomedical literature. Conventional text mining pipelines provide tools to automatically extract entities and relationships from the scientific literature, and partially replace expert curation, but such machine learning frameworks often require a large amount of labeled training data and thus lack scalability for both larger document corpora and new relationship types. RESULTS: We developed an application of Snorkel, a weakly supervised learning framework, for extracting chemical reaction relationships from biomedical literature abstracts. For this work, we defined a chemical reaction relationship as the transformation of chemical A to chemical B. We built and evaluated our system on small annotated sets of chemical reaction relationships from two corpora: curated bacteria-related abstracts from the MetaCyc database (MetaCyc_Corpus) and a more general set of abstracts annotated with MeSH (Medical Subject Headings) term Bacteria (Bacteria_Corpus; a superset of MetaCyc_Corpus). For the MetaCyc_Corpus, we obtained 84% precision and 41% recall (55% F1 score). Extending to the more general Bacteria_Corpus decreased precision to 62% with only a four-point drop in recall to 37% (46% F1 score). Overall, the Bacteria_Corpus contained two orders of magnitude more candidate chemical reaction relationships (nine million candidates vs 68,0000 candidates) and had a larger class imbalance (2.5% positives vs 5% positives) as compared to the MetaCyc_Corpus. In total, we extracted 6871 chemical reaction relationships from nine million candidates in the Bacteria_Corpus. CONCLUSIONS: With this work, we built a database of chemical reaction relationships from almost 900,000 scientific abstracts without a large training set of labeled annotations. Further, we showed the generalizability of our initial application built on MetaCyc documents enriched with chemical reactions to a general set of articles related to bacteria. Emily K. Mallory, Matthieu de Rochemonteix, Alexander Ratner, Ambika Acharya, Christopher Ré, Roselie A. Bright, Russ B. Altman |
BMC Bioinform. | 7 |
| 2020 | OrderRex clinical user testing: a randomized trial of recommender system decision support on simulated casesabstractOBJECTIVE: To assess usability and usefulness of a machine learning-based order recommender system applied to simulated clinical cases. MATERIALS AND METHODS: 43 physicians entered orders for 5 simulated clinical cases using a clinical order entry interface with or without access to a previously developed automated order recommender system. Cases were randomly allocated to the recommender system in a 3:2 ratio. A panel of clinicians scored whether the orders placed were clinically appropriate. Our primary outcome included the difference in clinical appropriateness scores. Secondary outcomes included total number of orders, case time, and survey responses. RESULTS: Clinical appropriateness scores per order were comparable for cases randomized to the order recommender system (mean difference -0.11 order per score, 95% CI: [-0.41, 0.20]). Physicians using the recommender placed more orders (median 16 vs 15 orders, incidence rate ratio 1.09, 95%CI: [1.01-1.17]). Case times were comparable with the recommender system. Order suggestions generated from the recommender system were more likely to match physician needs than standard manual search options. Physicians used recommender suggestions in 98% of available cases. Approximately 95% of participants agreed the system would be useful for their workflows. DISCUSSION: User testing with a simulated electronic medical record interface can assess the value of machine learning and clinical decision support tools for clinician usability and acceptance before live deployments. CONCLUSIONS: Clinicians can use and accept machine learned clinical order recommendations integrated into an electronic order entry interface in a simulated setting. The clinical appropriateness of orders entered was comparable even when supported by automated recommendations. Andre Kumar, Rachael C. Aikens, Jason Hom, Lisa Shieh, Jonathan Chiang, David Morales, Divya Saini, Mark A. Musen, Michael T. M. Baiocchi, Russ B. Altman, Mary K. Goldstein, Steven M. Asch, Jonathan H. Chen |
J. Am. Medical Informatics Assoc. | 10 |
| 2020 | Classifying non-small cell lung cancer types and transcriptomic subtypes using convolutional neural networksabstractOBJECTIVE: Non-small cell lung cancer is a leading cause of cancer death worldwide, and histopathological evaluation plays the primary role in its diagnosis. However, the morphological patterns associated with the molecular subtypes have not been systematically studied. To bridge this gap, we developed a quantitative histopathology analytic framework to identify the types and gene expression subtypes of non-small cell lung cancer objectively. MATERIALS AND METHODS: We processed whole-slide histopathology images of lung adenocarcinoma (n = 427) and lung squamous cell carcinoma patients (n = 457) in the Cancer Genome Atlas. We built convolutional neural networks to classify histopathology images, evaluated their performance by the areas under the receiver-operating characteristic curves (AUCs), and validated the results in an independent cohort (n = 125). RESULTS: To establish neural networks for quantitative image analyses, we first built convolutional neural network models to identify tumor regions from adjacent dense benign tissues (AUCs > 0.935) and recapitulated expert pathologists' diagnosis (AUCs > 0.877), with the results validated in an independent cohort (AUCs = 0.726-0.864). We further demonstrated that quantitative histopathology morphology features identified the major transcriptomic subtypes of both adenocarcinoma and squamous cell carcinoma (P < .01). DISCUSSION: Our study is the first to classify the transcriptomic subtypes of non-small cell lung cancer using fully automated machine learning methods. Our approach does not rely on prior pathology knowledge and can discover novel clinically relevant histopathology patterns objectively. The developed procedure is generalizable to other tumor types or diseases. Kun-Hsing Yu, Gerald J. Berry, Christopher Ré, Russ B. Altman, Michael Snyder 0001, Isaac S. Kohane |
J. Am. Medical Informatics Assoc. | 5 |
| 2020 | Transfer learning enables prediction of CYP2D6 haplotype functionabstractCytochrome P450 2D6 (CYP2D6) is a highly polymorphic gene whose protein product metabolizes more than 20% of clinically used drugs. Genetic variations in CYP2D6 are responsible for interindividual heterogeneity in drug response that can lead to drug toxicity and ineffective treatment, making CYP2D6 one of the most important pharmacogenes. Prediction of CYP2D6 phenotype relies on curation of literature-derived functional studies to assign a functional status to CYP2D6 haplotypes. As the number of large-scale sequencing efforts grows, new haplotypes continue to be discovered, and assignment of function is challenging to maintain. To address this challenge, we have trained a convolutional neural network to predict functional status of CYP2D6 haplotypes, called Hubble.2D6. Hubble.2D6 predicts haplotype function from sequence data and was trained using two pre-training steps with a combination of real and simulated data. We find that Hubble.2D6 predicts CYP2D6 haplotype functional status with 88% accuracy in a held-out test set and explains 47.5% of the variance in in vitro functional data among star alleles with unknown function. Hubble.2D6 may be a useful tool for assigning function to haplotypes with uncurated function, and used for screening individuals who are at risk of being poor metabolizers. Gregory McInnes, Rachel Dalton, Katrin Sangkuhl, Michelle Whirl Carrillo, Seung-been Lee, Philip S. Tsao, Andrea Gaedigk, Russ B. Altman, Erica L. Woodahl |
PLoS Comput. Biol. | 8 |
| 2019 | Classifying Non-Small Cell Lung Cancer Histopathology Types and Transcriptomic Subtypes using Convolutional Neural Networks
Kun-Hsing Yu, Gerald J. Berry, Christopher Ré, Russ B. Altman, Michael Snyder 0001, Isaac S. Kohane |
AMIA | 5 |
| 2019 | GRep: Gene Set Representation via Gaussian Embedding
Sheng Wang 0012, Emily R. Flynn, Russ B. Altman |
RECOMB | 3 |
| 2019 | Computational analysis of kinase inhibitor selectivity using structural knowledgeabstractMotivation: Kinases play a significant role in diverse disease signaling pathways and understanding kinase inhibitor selectivity, the tendency of drugs to bind to off-targets, remains a top priority for kinase inhibitor design and clinical safety assessment. Traditional approaches for kinase selectivity analysis using biochemical activity and binding assays are useful but can be costly and are often limited by the kinases that are available. On the other hand, current computational kinase selectivity prediction methods are computational intensive and can rarely achieve sufficient accuracy for large-scale kinome wide inhibitor selectivity profiling. Results: Here, we present a KinomeFEATURE database for kinase binding site similarity search by comparing protein microenvironments characterized using diverse physiochemical descriptors. Initial selectivity prediction of 15 known kinase inhibitors achieved an >90% accuracy and demonstrated improved performance in comparison to commonly used kinase inhibitor selectivity prediction methods. Additional kinase ATP binding site similarity assessment (120 binding sites) identified 55 kinases with significant promiscuity and revealed unexpected inhibitor cross-activities between PKR and FGFR2 kinases. Kinome-wide selectivity profiling of 11 kinase drug candidates predicted novel as well as experimentally validated off-targets and suggested structural mechanisms of kinase cross-activities. Our study demonstrated potential utilities of our approach for large-scale kinase inhibitor selectivity profiling that could contribute to kinase drug development and safety assessment. Availability and implementation: The KinomeFEATURE database and the associated scripts for performing kinase pocket similarity search can be downloaded from the Stanford SimTK website (https://simtk.org/projects/kdb). Supplementary information: Supplementary data are available at Bioinformatics online. Yu-Chen Lo, Tianyun Liu, Kari M. Morrissey, Satoko Kakiuchi-Kiyota, Adam R. Johnson, Fabio Broccatelli, Amita Joshi, Russ B. Altman |
Bioinform. | 9 |
| 2019 | High precision protein functional site detection using 3D convolutional neural networksabstractMOTIVATION: Accurate annotation of protein functions is fundamental for understanding molecular and cellular physiology. Data-driven methods hold promise for systematically deriving rules underlying the relationship between protein structure and function. However, the choice of protein structural representation is critical. Pre-defined biochemical features emphasize certain aspects of protein properties while ignoring others, and therefore may fail to capture critical information in complex protein sites. RESULTS: In this paper, we present a general framework that applies 3D convolutional neural networks (3DCNNs) to structure-based protein functional site detection. The framework can extract task-dependent features automatically from the raw atom distributions. We benchmarked our method against other methods and demonstrate better or comparable performance for site detection. Our deep 3DCNNs achieved an average recall of 0.955 at a precision threshold of 0.99 on PROSITE families, detected 98.89 and 92.88% of nitric oxide synthase and TRYPSIN-like enzyme sites in Catalytic Site Atlas, and showed good performance on challenging cases where sequence motifs are absent but a function is known to exist. Finally, we inspected the individual contributions of each atom to the classification decisions and show that our models successfully recapitulate known 3D features within protein functional sites. AVAILABILITY AND IMPLEMENTATION: The 3DCNN models described in this paper are available at https://simtk.org/projects/fscnn. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Wen Torng, Russ B. Altman |
Bioinform. | 2 |
| 2019 | PathFXweb: a web application for identifying drug safety and efficacy phenotypesabstractSUMMARY: Limited efficacy and intolerable safety limit therapeutic development and identification of potential liabilities earlier in development could significantly improve this process. Computational approaches which aggregate data from multiple sources and consider the drug's pathways effects could add to identification of these liabilities earlier. Such computational methods must be accessible to a variety of users beyond computational scientists, especially regulators and industry scientists, in order to impact the therapeutic development process. We have previously developed and published PathFX, an algorithm for identifying drug networks and phenotypes for understanding drug associations to safety and efficacy. Here we present a streamlined and easy-to-use PathFX web application that allows users to search for drug networks and associated phenotypes. We have also added visualization, and phenotype clustering to improve functionality and interpretability of PathFXweb. AVAILABILITY AND IMPLEMENTATION: https://www.pathfxweb.net/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jennifer L. Wilson, Mike Wong 0001, Ajinkya Chalke, Nicholas Stepanov, Dragutin Petkovic, Russ B. Altman |
Bioinform. | 6 |
| 2019 | RedMed: Extending drug lexicons for social media applications
Adam Lavertu, Russ B. Altman |
J. Biomed. Informatics | 2 |
| 2018 | Unraveling the Molecular Basis of Lung Adenocarcinoma Dedifferentiation and Prognosis by Integrating Omics and Histopathology
Kun-Hsing Yu, Gerald J. Berry, Daniel L. Rubin, Christopher Ré, Russ B. Altman, Michael Snyder 0001 |
AMIA | 5 |
| 2018 | A probabilistic pathway score (PROPS) for classification with applications to inflammatory bowel diseaseabstractSummary: Gene-based supervised machine learning classification models have been widely used to differentiate disease states, predict disease progression and determine effective treatment options. However, many of these classifiers are sensitive to noise and frequently do not replicate in external validation sets. For complex, heterogeneous diseases, these classifiers are further limited by being unable to capture varying combinations of genes that lead to the same phenotype. Pathway-based classification can overcome these challenges by using robust, aggregate features to represent biological mechanisms. In this work, we developed a novel pathway-based approach, PRObabilistic Pathway Score, which uses genes to calculate individualized pathway scores for classification. Unlike previous individualized pathway-based classification methods that use gene sets, we incorporate gene interactions using probabilistic graphical models to more accurately represent the underlying biology and achieve better performance. We apply our method to differentiate two similar complex diseases, ulcerative colitis (UC) and Crohn's disease (CD), which are the two main types of inflammatory bowel disease (IBD). Using five IBD datasets, we compare our method against four gene-based and four alternative pathway-based classifiers in distinguishing CD from UC. We demonstrate superior classification performance and provide biological insight into the top pathways separating CD from UC. Availability and Implementation: PROPS is available as a R package, which can be downloaded at http://simtk.org/home/props or on Bioconductor. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Lichy Han, Mateusz Maciejewski, Christoph Brockel, William Gordon, Scott B. Snapper, Joshua R. Korzenik, Lovisa Afzelius, Russ B. Altman |
Bioinform. | 8 |
| 2018 | A global network of biomedical relationships derived from textabstractMotivation: The biomedical community's collective understanding of how chemicals, genes and phenotypes interact is distributed across the text of over 24 million research articles. These interactions offer insights into the mechanisms behind higher order biochemical phenomena, such as drug-drug interactions and variations in drug response across individuals. To assist their curation at scale, we must understand what relationship types are possible and map unstructured natural language descriptions onto these structured classes. We used NCBI's PubTator annotations to identify instances of chemical, gene and disease names in Medline abstracts and applied the Stanford dependency parser to find connecting dependency paths between pairs of entities in single sentences. We combined a published ensemble biclustering algorithm (EBC) with hierarchical clustering to group the dependency paths into semantically-related categories, which we annotated with labels, or 'themes' ('inhibition' and 'activation', for example). We evaluated our theme assignments against six human-curated databases: DrugBank, Reactome, SIDER, the Therapeutic Target Database, OMIM and PharmGKB. Results: Clustering revealed 10 broad themes for chemical-gene relationships, 7 for chemical-disease, 10 for gene-disease and 9 for gene-gene. In most cases, enriched themes corresponded directly to known database relationships. Our final dataset, represented as a network, contained 37 491 thematically-labeled chemical-gene edges, 2 021 192 chemical-disease edges, 136 206 gene-disease edges and 41 418 gene-gene edges, each representing a single-sentence description of an interaction from somewhere in the literature. Availability and implementation: The complete network is available on Zenodo (https://zenodo.org/record/1035500). We have also provided the full set of dependency paths connecting biomedical entities in Medline abstracts, with associated sentences, for future use by the biomedical research community. Supplementary information: Supplementary data are available at Bioinformatics online. Bethany Percha, Russ B. Altman |
Bioinform. | 2 |
| 2018 | Data-driven human transcriptomic modules determined by independent component analysisabstractBACKGROUND: Analyzing the human transcriptome is crucial in advancing precision medicine, and the plethora of over half a million human microarray samples in the Gene Expression Omnibus (GEO) has enabled us to better characterize biological processes at the molecular level. However, transcriptomic analysis is challenging because the data is inherently noisy and high-dimensional. Gene set analysis is currently widely used to alleviate the issue of high dimensionality, but the user-defined choice of gene sets can introduce biasness in results. In this paper, we advocate the use of a fixed set of transcriptomic modules for such analysis. We apply independent component analysis to the large collection of microarray data in GEO in order to discover reproducible transcriptomic modules that can be used as features for machine learning. We evaluate the usability of these modules across six studies, and demonstrate (1) their usage as features for sample classification, and also their robustness in dealing with small training sets, (2) their regularization of data when clustering samples and (3) the biological relevancy of differentially expressed features. RESULTS: We identified 139 reproducible transcriptomic modules, which we term fundamental components (FCs). In studies with less than 50 samples, FC-space classification model outperformed their gene-space counterparts, with higher sensitivity (p < 0.01). The models also had higher accuracy and negative predictive value (p < 0.01) for small data sets (less than 30 samples). Additionally, we observed a reduction in batch effects when data is clustered in the FC-space. Finally, we found that differentially expressed FCs mapped to GO terms that were also identified via traditional gene-based approaches. CONCLUSIONS: The 139 FCs provide biologically-relevant summarization of transcriptomic data, and their performance in low sample settings suggest that they should be employed in such studies in order to harness the data efficiently. Weizhuang Zhou, Russ B. Altman |
BMC Bioinform. | 2 |
| 2018 | Expanding a radiology lexicon using contextual patterns in radiology reportsabstractObjective: Distributional semantics algorithms, which learn vector space representations of words and phrases from large corpora, identify related terms based on contextual usage patterns. We hypothesize that distributional semantics can speed up lexicon expansion in a clinical domain, radiology, by unearthing synonyms from the corpus. Materials and Methods: We apply word2vec, a distributional semantics software package, to the text of radiology notes to identify synonyms for RadLex, a structured lexicon of radiology terms. We stratify performance by term category, term frequency, number of tokens in the term, vector magnitude, and the context window used in vector building. Results: Ranking candidates based on distributional similarity to a target term results in high curation efficiency: on a ranked list of 775 249 terms, >50% of synonyms occurred within the first 25 terms. Synonyms are easier to find if the target term is a phrase rather than a single word, if it occurs at least 100× in the corpus, and if its vector magnitude is between 4 and 5. Some RadLex categories, such as anatomical substances, are easier to identify synonyms for than others. Discussion: The unstructured text of clinical notes contains a wealth of information about human diseases and treatment patterns. However, searching and retrieving information from clinical notes often suffer due to variations in how similar concepts are described in the text. Biomedical lexicons address this challenge, but are expensive to produce and maintain. Distributional semantics algorithms can assist lexicon curation, saving researchers time and money. Bethany Percha, Yuhao Zhang 0004, Selen Bozkurt, Daniel L. Rubin, Russ B. Altman, Curt Langlotz |
J. Am. Medical Informatics Assoc. | 5 |
| 2018 | PathFX provides mechanistic insights into drug efficacy and safety for regulatory review and therapeutic developmentabstractFailure to demonstrate efficacy and safety issues are important reasons that drugs do not reach the market. An incomplete understanding of how drugs exert their effects hinders regulatory and pharmaceutical industry projections of a drug's benefits and risks. Signaling pathways mediate drug response and while many signaling molecules have been characterized for their contribution to disease or their role in drug side effects, our knowledge of these pathways is incomplete. To better understand all signaling molecules involved in drug response and the phenotype associations of these molecules, we created a novel method, PathFX, a non-commercial entity, to identify these pathways and drug-related phenotypes. We benchmarked PathFX by identifying drugs' marketed disease indications and reported a sensitivity of 41%, a 2.7-fold improvement over similar approaches. We then used PathFX to strengthen signals for drug-adverse event pairs occurring in the FDA Adverse Event Reporting System (FAERS) and also identified opportunities for drug repurposing for new diseases based on interaction paths that associated a marketed drug to that disease. By discovering molecular interaction pathways, PathFX improved our understanding of drug associations to safety and efficacy phenotypes. The algorithm may provide a new means to improve regulatory and therapeutic development decisions. Jennifer L. Wilson, Rebecca Racz, Tianyun Liu, Oluseyi Adeniyi, Jielin Sun, Anuradha Ramamoorthy, Michael Pacanowski, Russ B. Altman |
PLoS Comput. Biol. | 8 |
| 2017 | Predicting Non-Small Cell Lung Cancer Diagnosis and Prognosis by Fully Automated Microscopic Pathology Image Features
Kun-Hsing Yu, Ce Zhang 0001, Gerald J. Berry, Russ B. Altman, Christopher Ré, Daniel L. Rubin, Michael Snyder 0001 |
AMIA | 4 |
| 2017 | Imputing gene expression to maximize platform compatibilityabstractMicroarray measurements of gene expression constitute a large fraction of publicly shared biological data, and are available in the Gene Expression Omnibus (GEO). Many studies use GEO data to shape hypotheses and improve statistical power. Within GEO, the Affymetrix HG-U133A and HG-U133 Plus 2.0 are the two most commonly used microarray platforms for human samples; the HG-U133 Plus 2.0 platform contains 54 220 probes and the HG-U133A array contains a proper subset (21 722 probes). When different platforms are involved, the subset of common genes is most easily compared. This approach results in the exclusion of substantial measured data and can limit downstream analysis. To predict the expression values for the genes unique to the HG-U133 Plus 2.0 platform, we constructed a series of gene expression inference models based on genes common to both platforms. Our model predicts gene expression values that are within the variability observed in controlled replicate studies and are highly correlated with measured data. Using six previously published studies, we also demonstrate the improved performance of the enlarged feature space generated by our model in downstream analysis. Availability and Implementation: The gene inference model described in this paper is available as a R package (affyImpute), which can be downloaded at http://simtk.org/home/affyimpute. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Weizhuang Zhou, Lichy Han, Russ B. Altman |
Bioinform. | 3 |
| 2017 | 3D deep convolutional neural networks for amino acid environment similarity analysisabstractBACKGROUND: Central to protein biology is the understanding of how structural elements give rise to observed function. The surfeit of protein structural data enables development of computational methods to systematically derive rules governing structural-functional relationships. However, performance of these methods depends critically on the choice of protein structural representation. Most current methods rely on features that are manually selected based on knowledge about protein structures. These are often general-purpose but not optimized for the specific application of interest. In this paper, we present a general framework that applies 3D convolutional neural network (3DCNN) technology to structure-based protein analysis. The framework automatically extracts task-specific features from the raw atom distribution, driven by supervised labels. As a pilot study, we use our network to analyze local protein microenvironments surrounding the 20 amino acids, and predict the amino acids most compatible with environments within a protein structure. To further validate the power of our method, we construct two amino acid substitution matrices from the prediction statistics and use them to predict effects of mutations in T4 lysozyme structures. RESULTS: Our deep 3DCNN achieves a two-fold increase in prediction accuracy compared to models that employ conventional hand-engineered features and successfully recapitulates known information about similar and different microenvironments. Models built from our predictions and substitution matrices achieve an 85% accuracy predicting outcomes of the T4 lysozyme mutation variants. Our substitution matrices contain rich information relevant to mutation analysis compared to well-established substitution matrices. Finally, we present a visualization method to inspect the individual contributions of each atom to the classification decisions. CONCLUSIONS: End-to-end trained deep learning networks consistently outperform methods using hand-engineered features, suggesting that the 3DCNN framework is well suited for analysis of protein microenvironments and may be useful for other protein structural analyses. Wen Torng, Russ B. Altman |
BMC Bioinform. | 2 |
| 2017 | Predicting inpatient clinical order patterns with probabilistic topic models vs conventional order setsabstractOBJECTIVE: Build probabilistic topic model representations of hospital admissions processes and compare the ability of such models to predict clinical order patterns as compared to preconstructed order sets. MATERIALS AND METHODS: The authors evaluated the first 24 hours of structured electronic health record data for > 10 K inpatients. Drawing an analogy between structured items (e.g., clinical orders) to words in a text document, the authors performed latent Dirichlet allocation probabilistic topic modeling. These topic models use initial clinical information to predict clinical orders for a separate validation set of > 4 K patients. The authors evaluated these topic model-based predictions vs existing human-authored order sets by area under the receiver operating characteristic curve, precision, and recall for subsequent clinical orders. RESULTS: Existing order sets predict clinical orders used within 24 hours with area under the receiver operating characteristic curve 0.81, precision 16%, and recall 35%. This can be improved to 0.90, 24%, and 47% ( P < 10 -20 ) by using probabilistic topic models to summarize clinical data into up to 32 topics. Many of these latent topics yield natural clinical interpretations (e.g., "critical care," "pneumonia," "neurologic evaluation"). DISCUSSION: Existing order sets tend to provide nonspecific, process-oriented aid, with usability limitations impairing more precise, patient-focused support. Algorithmic summarization has the potential to breach this usability barrier by automatically inferring patient context, but with potential tradeoffs in interpretability. CONCLUSION: Probabilistic topic modeling provides an automated approach to detect thematic trends in patient care and generate decision support content. A potential use case finds related clinical orders for decision support. Jonathan H. Chen, Mary K. Goldstein, Steven M. Asch, Lester Mackey, Russ B. Altman |
J. Am. Medical Informatics Assoc. | 5 |
| 2017 | Development of an automated assessment tool for MedWatch reports in the FDA adverse event reporting systemabstractOBJECTIVE: As the US Food and Drug Administration (FDA) receives over a million adverse event reports associated with medication use every year, a system is needed to aid FDA safety evaluators in identifying reports most likely to demonstrate causal relationships to the suspect medications. We combined text mining with machine learning to construct and evaluate such a system to identify medication-related adverse event reports. METHODS: FDA safety evaluators assessed 326 reports for medication-related causality. We engineered features from these reports and constructed random forest, L1 regularized logistic regression, and support vector machine models. We evaluated model accuracy and further assessed utility by generating report rankings that represented a prioritized report review process. RESULTS: Our random forest model showed the best performance in report ranking and accuracy, with an area under the receiver operating characteristic curve of 0.66. The generated report ordering assigns reports with a higher probability of medication-related causality a higher rank and is significantly correlated to a perfect report ordering, with a Kendall's tau of 0.24 ( P = .002). CONCLUSION: Our models produced prioritized report orderings that enable FDA safety evaluators to focus on reports that are more likely to contain valuable medication-related adverse event information. Applying our models to all FDA adverse event reports has the potential to streamline the manual review process and greatly reduce reviewer workload. Lichy Han, Robert Ball, Carol A. Pamer, Russ B. Altman, Scott Proestel |
J. Am. Medical Informatics Assoc. | 4 |
| 2016 | Usability of an Automated Recommender System for Clinical Order Entry
Jonathan H. Chen, Mary K. Goldstein, Steven M. Asch, Russ B. Altman |
AMIA | 4 |
| 2016 | Current Progress in Bioinformatics 2016abstractIn this issue, we present five review articles focusing on active and emerging areas of bioinformatics. We identified these areas as ones in which there has recently been a critical mass of initial publications that set the direction of the field, and lay out the key scientific challenges going forward. Much of current activity in bioinformatics is informed and inspired by the recent increased interest in ‘precision medicine' that US President Barack Obama highlighted in his January 2015 State of the Union address. It is clear to all that the computational techniques will be mandatory in the design and delivery of precise healthcare. Li et al. provide an overview of the progress in computational methods for drug repositioning. The cost of drug development is very high, and once approved a drug can be used by healthcare providers for ‘off label' indications. Although most drugs are approved based on their safety and efficacy in the context of a limited set of disease indications, it is also possible that they have salutary effects for other diseases. Indeed, the unwanted ‘side effects' of a drug may be an important clue to its utility for other diseases. In this review, the authors summarize the key data sources available for computational approaches, review the computational techniques used to associate drugs with new indications and provide a perspective on the most compelling current use cases. Success in drug repositioning leverages the huge investment in drugs that are on the market, and also open up opportunities for drug combinations that are useful. Tyler et al. provide a useful review on our current understanding of pleiotropy—the association of a single genetic locus with multiple phenotypes. There has been great interest in pleiotropy as an indicator of potential shared genetic architecture between phenotypes that may previously have been considered unrelated. The discovery of relations can suggest new molecular mechanisms, clinical correlations and potentially unexpected drug response (because of the role of a target in multiple phenotypes). This review focuses on computational methods to discover and characterize pleiotropy, and examines how these characterizations are useful for understanding the underlying genetic architecture and opportunities for novel discoveries. Khare et al. summarize the current interest in using crowdsourcing in biomedical research. Crowdsourcing generally refers to the activity of engaging large numbers of people who make efforts on behalf of science, often without formal scientific credentials. The authors define two groups of users: those who provide data based on their behavior (e.g. search logs, Facebook posts, tweets) that can be used for discovery or hypothesis generation, and users who actively provide labor to support scientific research. Each of these user types creates challenges and opportunities for scientists (while raising interesting issues of authorship and human subjects research). This review summarizes recent work in this relatively recent and fascinating phenomenon. Gonzalez et al. provide a useful review of text and data mining for biomedical discovery, particularly in the context of precision medicine. After defining the basic concepts and methods in text mining, they review recent emerging applications including the extraction of molecular pathways, the prediction of gene function, drug repositioning, data integration and pharmacogenomics. As long as our scientific colleagues insist on reporting their results in natural language, there will be a challenge of computational analysis of text, and the creation of structured databases of information reported using natural language. Finally, Greene et al. provide a review of the competencies required for professionals working in biomedical data science or ‘big data' in biomedicine. Existing curricula focusing on computational biology and biomedical informatics have recently been challenged to expand and augment to accommodate the great academic and industrial interest in biomedical data science. The authors discuss the typical differences between modern data science curricular needs, and those present in more traditional programs. Notably, they suggest additional courses that may augment existing curricula, including in ‘biological information flow', ‘statistical challenges of big data' and ‘computational challenges of big data'. We hope you enjoy these reviews of important emerging areas, and join us in marveling at the fantastic opportunities and challenges that continue to provide a rich scientific research agenda for bioinformatics. Russ B. Altman |
Briefings Bioinform. | 1 |
| 2016 | STAMS: STRING-assisted module search for genome wide association studies and application to autismabstractMotivation: Analyzing genome wide association data in the context of biological pathways helps us understand how genetic variation influences phenotype and increases power to find associations. However, the utility of pathway-based analysis tools is hampered by undercuration and reliance on a distribution of signal across all of the genes in a pathway. Methods that combine genome wide association results with genetic networks to infer the key phenotype-modulating subnetworks combat these issues, but have primarily been limited to network definitions with yes/no labels for gene-gene interactions. A recent method (EW_dmGWAS) incorporates a biological network with weighted edge probability by requiring a secondary phenotype-specific expression dataset. In this article, we combine an algorithm for weighted-edge module searching and a probabilistic interaction network in order to develop a method, STAMS, for recovering modules of genes with strong associations to the phenotype and probable biologic coherence. Our method builds on EW_dmGWAS but does not require a secondary expression dataset and performs better in six test cases. Results: We show that our algorithm improves over EW_dmGWAS and standard gene-based analysis by measuring precision and recall of each method on separately identified associations. In the Wellcome Trust Rheumatoid Arthritis study, STAMS-identified modules were more enriched for separately identified associations than EW_dmGWAS (STAMS P-value 3.0 × 10−4; EW_dmGWAS- P-value = 0.8). We demonstrate that the area under the Precision-Recall curve is 5.9 times higher with STAMS than EW_dmGWAS run on the Wellcome Trust Type 1 Diabetes data. Availability and Implementation: STAMS is implemented as an R package and is freely available at https://simtk.org/projects/stams. Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Sara Hillenmeyer, Lea K. Davis, Eric R. Gamazon, Edwin H. Cook Jr., Nancy J. Cox, Russ B. Altman |
Bioinform. | 6 |
| 2016 | Large-scale extraction of gene interactions from full-text literature using DeepDiveabstractMOTIVATION: A complete repository of gene-gene interactions is key for understanding cellular processes, human disease and drug response. These gene-gene interactions include both protein-protein interactions and transcription factor interactions. The majority of known interactions are found in the biomedical literature. Interaction databases, such as BioGRID and ChEA, annotate these gene-gene interactions; however, curation becomes difficult as the literature grows exponentially. DeepDive is a trained system for extracting information from a variety of sources, including text. In this work, we used DeepDive to extract both protein-protein and transcription factor interactions from over 100,000 full-text PLOS articles. METHODS: We built an extractor for gene-gene interactions that identified candidate gene-gene relations within an input sentence. For each candidate relation, DeepDive computed a probability that the relation was a correct interaction. We evaluated this system against the Database of Interacting Proteins and against randomly curated extractions. RESULTS: Our system achieved 76% precision and 49% recall in extracting direct and indirect interactions involving gene symbols co-occurring in a sentence. For randomly curated extractions, the system achieved between 62% and 83% precision based on direct or indirect interactions, as well as sentence-level and document-level precision. Overall, our system extracted 3356 unique gene pairs using 724 features from over 100,000 full-text articles. AVAILABILITY AND IMPLEMENTATION: Application source code is publicly available at https://github.com/edoughty/deepdive_genegene_app CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Emily K. Mallory, Ce Zhang 0001, Christopher Ré, Russ B. Altman |
Bioinform. | 4 |
| 2016 | OrderRex: clinical order decision support and outcome predictions by data-mining electronic medical recordsabstractOBJECTIVE: To answer a "grand challenge" in clinical decision support, the authors produced a recommender system that automatically data-mines inpatient decision support from electronic medical records (EMR), analogous to Netflix or Amazon.com's product recommender. MATERIALS AND METHODS: EMR data were extracted from 1 year of hospitalizations (>18K patients with >5.4M structured items including clinical orders, lab results, and diagnosis codes). Association statistics were counted for the ∼1.5K most common items to drive an order recommender. The authors assessed the recommender's ability to predict hospital admission orders and outcomes based on initial encounter data from separate validation patients. RESULTS: Compared to a reference benchmark of using the overall most common orders, the recommender using temporal relationships improves precision at 10 recommendations from 33% to 38% (P < 10(-10)) for hospital admission orders. Relative risk-based association methods improve inverse frequency weighted recall from 4% to 16% (P < 10(-16)). The framework yields a prediction receiver operating characteristic area under curve (c-statistic) of 0.84 for 30 day mortality, 0.84 for 1 week need for ICU life support, 0.80 for 1 week hospital discharge, and 0.68 for 30-day readmission. DISCUSSION: Recommender results quantitatively improve on reference benchmarks and qualitatively appear clinically reasonable. The method assumes that aggregate decision making converges appropriately, but ongoing evaluation is necessary to discern common behaviors from "correct" ones. CONCLUSIONS: Collaborative filtering recommender algorithms generate clinical decision support that is predictive of real practice patterns and clinical outcomes. Incorporating temporal relationships improves accuracy. Different evaluation metrics satisfy different goals (predicting likely events vs. "interesting" suggestions). Jonathan H. Chen, Tanya Podchiyska, Russ B. Altman |
J. Am. Medical Informatics Assoc. | 3 |
| 2016 | Computing disease incidence, prevalence and comorbidity from electronic medical records
Steven C. Bagley, Russ B. Altman |
J. Biomed. Informatics | 2 |
| 2016 | Constraints on Biological Mechanism from Disease Comorbidity Using Electronic Medical Records and Database of Genetic VariantsabstractPatterns of disease co-occurrence that deviate from statistical independence may represent important constraints on biological mechanism, which sometimes can be explained by shared genetics. In this work we study the relationship between disease co-occurrence and commonly shared genetic architecture of disease. Records of pairs of diseases were combined from two different electronic medical systems (Columbia, Stanford), and compared to a large database of published disease-associated genetic variants (VARIMED); data on 35 disorders were available across all three sources, which include medical records for over 1.2 million patients and variants from over 17,000 publications. Based on the sources in which they appeared, disease pairs were categorized as having predominant clinical, genetic, or both kinds of manifestations. Confounding effects of age on disease incidence were controlled for by only comparing diseases when they fall in the same cluster of similarly shaped incidence patterns. We find that disease pairs that are overrepresented in both electronic medical record systems and in VARIMED come from two main disease classes, autoimmune and neuropsychiatric. We furthermore identify specific genes that are shared within these disease groups. Steven C. Bagley, Marina Sirota, Richard O. Chen, Atul J. Butte, Russ B. Altman |
PLoS Comput. Biol. | 5 |
| 2015 | An ontology for Autism Spectrum Disorder (ASD) to infer ASD phenotypes from Autism Diagnostic Interview-Revised data
Omri Mugzach, Mor Peleg, Steven C. Bagley, Stephen J. Guter, Edwin H. Cook Jr., Russ B. Altman |
J. Biomed. Informatics | 6 |
| 2015 | Variations in the Binding Pocket of an Inhibitor of the Bacterial Division Protein FtsZ across Genotypes and SpeciesabstractThe recent increase in antibiotic resistance in pathogenic bacteria calls for new approaches to drug-target selection and drug development. Targeting the mechanisms of action of proteins involved in bacterial cell division bypasses problems associated with increasingly ineffective variants of older antibiotics; to this end, the essential bacterial cytoskeletal protein FtsZ is a promising target. Recent work on its allosteric inhibitor, PC190723, revealed in vitro activity on Staphylococcus aureus FtsZ and in vivo antimicrobial activities. However, the mechanism of drug action and its effect on FtsZ in other bacterial species are unclear. Here, we examine the structural environment of the PC190723 binding pocket using PocketFEATURE, a statistical method that scores the similarity between pairs of small-molecule binding sites based on 3D structure information about the local microenvironment, and molecular dynamics (MD) simulations. We observed that species and nucleotide-binding state have significant impacts on the structural properties of the binding site, with substantially disparate microenvironments for bacterial species not from the Staphylococcus genus. Based on PocketFEATURE analysis of MD simulations of S. aureus FtsZ bound to GTP or with mutations that are known to confer PC190723 resistance, we predict that PC190723 strongly prefers to bind Staphylococcus FtsZ in the nucleotide-bound state. Furthermore, MD simulations of an FtsZ dimer indicated that polymerization may enhance PC190723 binding. Taken together, our results demonstrate that a drug-binding pocket can vary significantly across species, genetic perturbations, and in different polymerization states, yielding important information for the further development of FtsZ inhibitors. Amanda Miguel, Jen Hsin, Tianyun Liu, Grace W. Tang, Russ B. Altman, Kerwyn Casey Huang |
PLoS Comput. Biol. | 5 |
| 2015 | Learning the Structure of Biomedical Relationships from Unstructured TextabstractThe published biomedical research literature encompasses most of our understanding of how drugs interact with gene products to produce physiological responses (phenotypes). Unfortunately, this information is distributed throughout the unstructured text of over 23 million articles. The creation of structured resources that catalog the relationships between drugs and genes would accelerate the translation of basic molecular knowledge into discoveries of genomic biomarkers for drug response and prediction of unexpected drug-drug interactions. Extracting these relationships from natural language sentences on such a large scale, however, requires text mining algorithms that can recognize when different-looking statements are expressing similar ideas. Here we describe a novel algorithm, Ensemble Biclustering for Classification (EBC), that learns the structure of biomedical relationships automatically from text, overcoming differences in word choice and sentence structure. We validate EBC's performance against manually-curated sets of (1) pharmacogenomic relationships from PharmGKB and (2) drug-target relationships from DrugBank, and use it to discover new drug-gene relationships for both knowledge bases. We then apply EBC to map the complete universe of drug-gene relationships based on their descriptions in Medline, revealing unexpected structure that challenges current notions about how these relationships are expressed in text. For instance, we learn that newer experimental findings are described in consistently different ways than established knowledge, and that seemingly pure classes of relationships can exhibit interesting chimeric structure. The EBC algorithm is flexible and adaptable to a wide range of problems in biomedical text mining. Bethany Percha, Russ B. Altman |
PLoS Comput. Biol. | 2 |
| 2014 | "Doctors who ordered this also ordered..." Automated physician order recommendations and outcome predictions by data-mining electronic medical records
Jonathan H. Chen, Russ B. Altman |
AMIA | 2 |
| 2014 | Environmental and State-Level Regulatory Factors Affect the Incidence of Autism and Intellectual DisabilityabstractMany factors affect the risks for neurodevelopmental maladies such as autism spectrum disorders (ASD) and intellectual disability (ID). To compare environmental, phenotypic, socioeconomic and state-policy factors in a unified geospatial framework, we analyzed the spatial incidence patterns of ASD and ID using an insurance claims dataset covering nearly one third of the US population. Following epidemiologic evidence, we used the rate of congenital malformations of the reproductive system as a surrogate for environmental exposure of parents to unmeasured developmental risk factors, including toxins. Adjusted for gender, ethnic, socioeconomic, and geopolitical factors, the ASD incidence rates were strongly linked to population-normalized rates of congenital malformations of the reproductive system in males (an increase in ASD incidence by 283% for every percent increase in incidence of malformations, 95% CI: [91%, 576%], p<6×10(-5)). Such congenital malformations were barely significant for ID (94% increase, 95% CI: [1%, 250%], p = 0.0384). Other congenital malformations in males (excluding those affecting the reproductive system) appeared to significantly affect both phenotypes: 31.8% ASD rate increase (CI: [12%, 52%], p<6×10(-5)), and 43% ID rate increase (CI: [23%, 67%], p<6×10(-5)). Furthermore, the state-mandated rigor of diagnosis of ASD by a pediatrician or clinician for consideration in the special education system was predictive of a considerable decrease in ASD and ID incidence rates (98.6%, CI: [28%, 99.99%], p = 0.02475 and 99% CI: [68%, 99.99%], p = 0.00637 respectively). Thus, the observed spatial variability of both ID and ASD rates is associated with environmental and state-level regulatory factors; the magnitude of influence of compound environmental predictors was approximately three times greater than that of state-level incentives. The estimated county-level random effects exhibited marked spatial clustering, strongly indicating existence of as yet unidentified localized factors driving apparent disease incidence. Finally, we found that the rates of ASD and ID at the county level were weakly but significantly correlated (Pearson product-moment correlation 0.0589, p = 0.00101), while for females the correlation was much stronger (0.197, p<2.26×10(-16)). Andrey Rzhetsky, Steven C. Bagley, Kanix Wang, Christopher S. Lyttle, Edwin H. Cook Jr., Russ B. Altman, Robert D. Gibbons |
PLoS Comput. Biol. | 6 |
| 2014 | Knowledge-based Fragment Binding PredictionabstractTarget-based drug discovery must assess many drug-like compounds for potential activity. Focusing on low-molecular-weight compounds (fragments) can dramatically reduce the chemical search space. However, approaches for determining protein-fragment interactions have limitations. Experimental assays are time-consuming, expensive, and not always applicable. At the same time, computational approaches using physics-based methods have limited accuracy. With increasing high-resolution structural data for protein-ligand complexes, there is now an opportunity for data-driven approaches to fragment binding prediction. We present FragFEATURE, a machine learning approach to predict small molecule fragments preferred by a target protein structure. We first create a knowledge base of protein structural environments annotated with the small molecule substructures they bind. These substructures have low-molecular weight and serve as a proxy for fragments. FragFEATURE then compares the structural environments within a target protein to those in the knowledge base to retrieve statistically preferred fragments. It merges information across diverse ligands with shared substructures to generate predictions. Our results demonstrate FragFEATURE's ability to rediscover fragments corresponding to the ligand bound with 74% precision and 82% recall on average. For many protein targets, it identifies high scoring fragments that are substructures of known inhibitors. FragFEATURE thus predicts fragments that can serve as inputs to fragment-based drug design or serve as refinement criteria for creating target-specific compound libraries for experimental or computational screening. Grace W. Tang, Russ B. Altman |
PLoS Comput. Biol. | 2 |
| 2013 | Expanding the Autism Ontology to DSM-IV Criteria
Omri Mugzach, Mor Peleg, Steven C. Bagley, Russ B. Altman |
AMIA | 4 |
| 2013 | Inferring the semantic relationships of words within an ontology using random indexing: applications to pharmacogenomics
Bethany Percha, Russ B. Altman |
AMIA | 2 |
| 2013 | Correspondence: Response to 'Use of an algorithm for identifying hidden drug-drug interactions in adverse event reports' by Gooden et alabstractCritical evaluation of the results of clinical studies is vital to the continued progress of medicine. We appreciate the work performed by Gooden and colleagues1 to evaluate the clinical significance of a drug interaction between paroxetine, a selective serotonin reuptake inhibitor, and pravastatin, a cholesterol-lowering statin, that we published previously.2 Our results demonstrated a 18.5 mg/dl increase in glucose levels in individuals without diabetes, and a 48 mg/dl increase in glucose level for diabetes patients using three electronic medical record systems. In the study, Gooden et al1 did not find a difference in the development of type 2 diabetes using administrative data. We agree that retrospective risk estimates such as ours may be influenced by selection biases, such as confounding by indication. However, in our replication and validation study3 we did not see increased glucose measurements for patients on other combinations of selective serotonin reuptake inhibitors and statins or for the two classes generally—patients who are expected to have the same comorbidities. We were also not able to identify any clinical reason for the existence of clinical confounders for this particular combination of drugs alone. Moreover, we note that prediabetic mice clearly showed a positive biological result and would not be subject to the same possible confounders as the human studies.3 The authors correctly point out that an increase in non-fasting blood glucose measurements may not lead to a clinically significant event, such as type 2 diabetes mellitus (T2DM). It is possible that the increase in random glucose is not sufficiently large result in a patient being newly diagnosed with diabetes. Moreover, our findings were for near-term changes in glucose; it is possible that over the longer term, glucose falls back to normal. This would require further investigation. Finally, patients with T2DM may have the disease for some time before a diagnosis is made. It is possible that the patients enrolled in the study by Gooden et al1 had not been observed long enough to note the development of diabetes if in fact such an observation does exist. To assess the clinical significance of the drug interaction Gooden et al1 evaluated the onset of new T2DM in all patients 18 years or older using claims data. Although administrative data constitute a powerful tool for evaluating disease, accrual of a single billing code for T2DM can falsely label patients as having diabetes (false positives) as well as also falsely excluding others as not having the disease (false negatives). For this reason, Ritchie et al4 and Kho et al5 both used phenotype algorithms for T2DM including laboratory values, medications, and diagnosis billing codes (also see PheKB.org). Using claims data alone may introduce too much noise and undermine the interpretation of the authors' analysis. Gooden et al1 correctly point out that non-fasting glucose values have high variance and are not uniformly collected for all patients. For this reason we performed a paired analysis that required a patient to have glucose laboratory tests run both before and after they began combination treatment with paroxetine and pravastatin.3 We found flat glucose measurements for the single-drug-only groups, which indicate that the variability in glucose laboratory tests is not enough to explain the divergence we see in patients on the combination.3 We fully agree with the authors closing sentiment that there should be careful separation of hypothesis generation (in our case an analysis of the US Food and Drug Administration's adverse event reporting system) and hypothesis testing (in our case replication in three electronic health record systems and validation in a mouse model). It is clear that evaluating the clinical significance of this interaction between these two commonly used drugs will require a deeper understanding of its mechanism, as well as the long-term consequences of exposure. None. Not commissioned; externally peer reviewed. Nicholas P. Tatonetti, Joshua C. Denny, Russ B. Altman |
J. Am. Medical Informatics Assoc. | 3 |
| 2013 | Web-scale pharmacovigilance: listening to signals from the crowdabstractAdverse drug events cause substantial morbidity and mortality and are often discovered after a drug comes to market. We hypothesized that Internet users may provide early clues about adverse drug events via their online information-seeking. We conducted a large-scale study of Web search log data gathered during 2010. We pay particular attention to the specific drug pairing of paroxetine and pravastatin, whose interaction was reported to cause hyperglycemia after the time period of the online logs used in the analysis. We also examine sets of drug pairs known to be associated with hyperglycemia and those not associated with hyperglycemia. We find that anonymized signals on drug interactions can be mined from search logs. Compared to analyses of other sources such as electronic health records (EHR), logs are inexpensive to collect and mine. The results demonstrate that logs of the search activities of populations of computer users can contribute to drug safety surveillance. Ryen W. White, Nicholas P. Tatonetti, Nigam H. Shah, Russ B. Altman, Eric Horvitz |
J. Am. Medical Informatics Assoc. | 4 |
| 2013 | K-Means for Parallel Architectures Using All-Prefix-Sum Sorting and Updating StepsabstractWe present an implementation of parallel K-means clustering, called Kps-means, that achieves high performance with near-full occupancy compute kernels without imposing limits on the number of dimensions and data points permitted as input, thus combining flexibility with high degrees of parallelism and efficiency. As a key element to performance improvement, we introduce parallel sorting as data preprocessing and updating steps. Our final implementation for Nvidia GPUs achieves speedups of up to 200-fold over CPU reference code and of up to three orders of magnitude when compared with popular numerical software packages. Kai Kohlhoff, Vijay S. Pande, Russ B. Altman |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2012 | Editorial: Current progress in Bioinformatics 2012abstractIn this issue, we present an annual review of progress in critical areas of biomedical computing and informatics. Our ability to collect and store biological and medical data has continued to increase at astounding rates over the last few years. As a result, the scientific and technical challenges for informatics and computing more generally have also increased. A colleague from a major drug company recently showed me a flyer with more than 20 job openings for informatics and computing. Graduates of programs in informatics are getting multiple offers (even in a sagging economy) in both industry and academia. There is a huge demand for the skill sets associated with the appropriate management and analysis of ‘big data’ in biology and medicine. In this issue, we provide a snapshot of some of the major areas of emerging effort in bioinformatics. As usual, we present the papers roughly in order of length scale from molecular to whole organism. Structural studies of large macromolecular complexes are increasingly conducted using hybrid models at several levels of resolution. In ‘Multiscale modeling of acromolecular biosystems’, Flores et al. summarize both the physics- and knowledge-based methods for representing the interactions between molecular components. Interestingly, ‘coarse graining’, the introduction of simplifications to improve computational complexity can be performed in either the length scale or the time scale. The authors also review recent progress in RNA modeling, docking and aggregation (relevant to many protein misfolding diseases). The applications of high-throughput molecular and cellular measurements have been protean and widespread. These data sources include next generation DNA sequencing, RNA expression analysis, metabolomics and varieties of proteomics. One of the prime applications is to understand the function (and dysfunction) of biological pathways, and it has become clear that the integration of multiple data sources leads to a clearer signal and minimizes the problems associated with systematic errors in any single measurement modality. In this review, Wang et al., ‘Identification of aberrant pathways and network activities from high-throughput data’, summarize recent work in modeling biological pathways and networks, and use high-throughput data to characterize biological dysfunction. Several interesting computational models have emerged, including some that capture spatial and temporal dynamics. It has become apparent in the last few years that microbial communities are ubiquitous on earth (including within healthy human hosts), and that they must be studied as a community, as the dependencies between individual species often stress our conventional ideas of species. Foster et al., in ‘Measuring the microbiome: perspectives on advances in DNA-based techniques for exploring microbial life’, describe how the arrival of next-generation sequencing capabilities has provided amazing new possibilities for accurately measuring the genomes present in microbial communities, while beginning to create the tools to dissect the relationships, functionalities and interactions between the molecular and cellular players. Two areas of technical progress recently have been statistical methods to characterize diversity and methods for visualization of the evolutionary and community relationships within microbial ecosystems. The emerging fields of systems biology and synthetic biology share a goal of creating accurate and predictive models of biological systems. One way to do this is to reverse engineer existing systems, in order to understand how they work. The review, ‘Reverse engineering bio-molecular systems (REBMS) using -omic data: challenges, progress and opportunities’ by Quo et al., lays out the challenges of integrating heterogeneous biochemical information, combining data mining with modeling approaches and creating experimental approaches towards validation of these reverse engineered systems. They review two case studies in the realm of breast cancer and bacterial chemotaxis, to show the power of this emerging toolkit. Translational medicine aims to bring the fruits of basic biological science into patient care. A major challenge for translational science is how to integrate high-throughput measurements and associate them with clinical entities, such as diseases, symptoms, drugs and diagnostic tests. One organizing principle for these data is ‘network biology’ in which the relationships between the entities are added to our understanding of the individual functions of the component genes. In this review, Bebek et al., ‘Network biology methods integrating biological data for translational science’, describe current progress in integrating human expression and interaction data with genome-wide data to understand disease. They describe applications in prioritization of candidate genes, building mechanistic models and applying these to particular models of disease. Recent advances in core natural language processing (NLP) technology and ontology-based annotation and markup have created exciting new possibilities for literature analysis. In particular, the application of these technologies to the problem of identifying, extracting and representing the relationships between genes and drugs has been active. Personalized medicine requires knowledge of how individual molecular entities function and how individual mutations can affect cellular and organ processes. The literature now contains well over 20 million citations, and so automated methods are mandatory for systematically surveying the literature. Hahn et al., in ‘Mining the pharmacogenomics literature–a survey of the state of the art’ discuss the fundamental activities of recognizing gene and drug names in text (named entity recognition), as well as the relationships between them. They also review infrastructural challenges and progress in creating gold-standard corpora with which to train and test new methodological advances. The arrival of next-generation sequencing has also enabled the marked acceleration of human genome sequencing, thus calling the question of how to interpret rare genetic variants in terms of disease risk and drug response. Capriotti et al., ‘Bioinformatics for personal genome interpretation’ first catalog the increasingly valuable and interconnected set of data resources that are now available. Next, they review methods for ranking candidate genes based on the likelihood that they play a role in some health-related phenotype. Finally, they discuss the emergence of new tools for analyzing novel variants within these genes, and end with a summary of the challenges to the field. We hope you enjoy this survey of important emerging areas in bioinformatics. Russ B. Altman |
Briefings Bioinform. | 1 |
| 2012 | Simbios: an NIH national center for physics-based simulation of biological structuresabstractPhysics-based simulation provides a powerful framework for understanding biological form and function. Simulations can be used by biologists to study macromolecular assemblies and by clinicians to design treatments for diseases. Simulations help biomedical researchers understand the physical constraints on biological systems as they engineer novel drugs, synthetic tissues, medical devices, and surgical interventions. Although individual biomedical investigators make outstanding contributions to physics-based simulation, the field has been fragmented. Applications are typically limited to a single physical scale, and individual investigators usually must create their own software. These conditions created a major barrier to advancing simulation capabilities. In 2004, we established a National Center for Physics-Based Simulation of Biological Structures (Simbios) to help integrate the field and accelerate biomedical research. In 6 years, Simbios has become a vibrant national center, with collaborators in 16 states and eight countries. Simbios focuses on problems at both the molecular scale and the organismal level, with a long-term goal of uniting these in accurate multiscale simulations. Scott L. Delp, Joy P. Ku, Vijay S. Pande, Michael A. Sherman, Russ B. Altman |
J. Am. Medical Informatics Assoc. | 5 |
| 2012 | A novel signal detection algorithm for identifying hidden drug-drug interactions in adverse event reportsabstractOBJECTIVE: Adverse drug events (ADEs) are common and account for 770 000 injuries and deaths each year and drug interactions account for as much as 30% of these ADEs. Spontaneous reporting systems routinely collect ADEs from patients on complex combinations of medications and provide an opportunity to discover unexpected drug interactions. Unfortunately, current algorithms for such "signal detection" are limited by underreporting of interactions that are not expected. We present a novel method to identify latent drug interaction signals in the case of underreporting. MATERIALS AND METHODS: We identified eight clinically significant adverse events. We used the FDA's Adverse Event Reporting System to build profiles for these adverse events based on the side effects of drugs known to produce them. We then looked for pairs of drugs that match these single-drug profiles in order to predict potential interactions. We evaluated these interactions in two independent data sets and also through a retrospective analysis of the Stanford Hospital electronic medical records. RESULTS: We identified 171 novel drug interactions (for eight adverse event categories) that are significantly enriched for known drug interactions (p=0.0009) and used the electronic medical record for independently testing drug interaction hypotheses using multivariate statistical models with covariates. CONCLUSION: Our method provides an option for detecting hidden interactions in spontaneous reporting systems by using side effect profiles to infer the presence of unreported adverse events. Nicholas P. Tatonetti, Guy Haskin Fernald, Russ B. Altman |
J. Am. Medical Informatics Assoc. | 3 |
| 2012 | The state of the art in text mining and natural language processing for pharmacogenomics
Adrien Coulet, Kevin Cohen 0001, Russ B. Altman |
J. Biomed. Informatics | 3 |
| 2012 | Introduction to Translational Bioinformatics CollectionabstractHow should we define translational bioinformatics? I had to answer this question unambiguously in March 2008 when I was asked to deliver a review of ‘‘recent progress in translational bioinformatics’’ at the American Medical Informatics Association’s Summit on Translational Bioinformatics. The lecture required me to define papers in the field, and then highlight exciting progress that occurred over the previous ,12 months. I have repeated this for the last few years, and the most difficult part of the exercise is limiting my review only to those papers that are within the field. I have never worried much about definitions within informatics fields; they tend to overlap, merge and evolve. ‘‘Informatics’’ seems clear: the study of how to represent, store, search, retrieve and analyze information. The adjectives in front of ‘‘informatics’’ vary but also tend to make sense: medical informatics concerns medical information, bioinformatics concerns basic biological information, clinical informatics focuses on the clinical delivery part of medical informatics, biomedical informatics merges bioinformatics and medical informatics, imaging informatics focuses on...images, and so on. So what does this adjective ‘‘translational’’ denote? Translational medical research has emerged as an important theme in the last decade. Starting with top-down leadership from the National Institutes of Health and its former Director, Dr. Elias Zerhouni, and moving through academic medical centers, research institutes and industrial research and development efforts, there has been interest in more effectively moving the discoveries and innovations in the laboratory to the bedside, leading to improved diagnosis, prognosis, and treatment. Translational research encompasses many activities including the creation of medical devices, molecular diagnostics, small molecule therapeutics, biological therapeutics, vaccines, and others. One of the main targets of translation, however, is revolutionary explosion of knowledge in molecular biology, genetics, and genomics. Some believe that the tremendous progress in discovery over the last 50+ years since elucidation of the double helix structure has not translated (there’s that word!) into much practical health benefit. While the accuracy of this claim can be debated, there can be no debate that our ability to measure (1) DNA sequence (including entire genomes!), (2) RNA sequence and expression, (3) protein sequence, structure, expression and modification, and (4) small molecule metabolite structure, presence, and quantity has advanced rapidly and enables us to imagine fantastic new technologies in pursuit of human health. There are many barriers to translating our molecular understanding into technologies that impact patients. These include understanding health market size and forces, the regulatory milieu, how to harden the technology for routine use, and how to navigate an increasingly complex intellectual property landscape. But before those activities can begin, we must overcome an even more fundamental barrier: connecting the stuff of molecular biology to the clinical world. Molecular and cellular biology studies genes, DNA, RNA messengers, microRNAs, proteins, signaling molecules and their cascades, metabolites, cellular communication processes and cellular organization. These data are freely available in valuable resources such as Genbank (http://www. ncbi.nlm.nih.gov/genbank/), Gene Expression Omnibus (http://www.ncbi.nlm. nih.gov/geo/), Protein Data Bank (http:// www.wwpdb.org/), KEGG (http://www. genome.jp/kegg/), MetaCyc (http:// metacyc.org/), Reactome (http://www. reactome.org), and many other resources. The clinical world studies diseases, signs, symptoms, drugs, patients, clinical laboratory measurements, and clinical images. The emergence of clinical and health information technologies has begun to make these clinical data available for research through biobanks, electronic medical records, FDA resources about drug labels and adverse events, and claims data. Therefore, a major challenge for translational medicine is to connect the molecular/cellular world with the clinical world. The published literature, available in PubMED (http://www.ncbi.nlm.nih. gov/pubmed), does this, as does the Unified Medical Language System (UMLS) that provides a lingua franca (http://www.nlm.nih.gov/research/umls/ ). However, it falls to translational bioinformatics to engineer the tools that link molecular/cellular entities and clinical entities. Thus, I define ‘‘translational bioinformatics’’ research as the development and application of informatics methods that connect molecular entities to clinical entities. In this collection, Dr. Kann and colleagues have assembled a wonderful group of authors to introduce the key threads of translational bioinformatics to those new to the field. The collection first provides Russ B. Altman |
PLoS Comput. Biol. | 1 |
| 2012 | Chapter 7: PharmacogenomicsabstractThere is great variation in drug-response phenotypes, and a "one size fits all" paradigm for drug delivery is flawed. Pharmacogenomics is the study of how human genetic information impacts drug response, and it aims to improve efficacy and reduced side effects. In this article, we provide an overview of pharmacogenetics, including pharmacokinetics (PK), pharmacodynamics (PD), gene and pathway interactions, and off-target effects. We describe methods for discovering genetic factors in drug response, including genome-wide association studies (GWAS), expression analysis, and other methods such as chemoinformatics and natural language processing (NLP). We cover the practical applications of pharmacogenomics both in the pharmaceutical industry and in a clinical setting. In drug discovery, pharmacogenomics can be used to aid lead identification, anticipate adverse events, and assist in drug repurposing efforts. Moreover, pharmacogenomic discoveries show promise as important elements of physician decision support. Finally, we consider the ethical, regulatory, and reimbursement challenges that remain for the clinical implementation of pharmacogenomics. Konrad J. Karczewski, Roxana Daneshjou, Russ B. Altman |
PLoS Comput. Biol. | 3 |
| 2011 | Bioinformatics challenges for personalized medicine
Guy Haskin Fernald, Emidio Capriotti, Roxana Daneshjou, Konrad J. Karczewski, Russ B. Altman |
Bioinform. | 5 |
| 2011 | Bioinformatics challenges for personalized medicineabstractMOTIVATION: Widespread availability of low-cost, full genome sequencing will introduce new challenges for bioinformatics. RESULTS: This review outlines recent developments in sequencing technologies and genome analysis methods for application in personalized medicine. New methods are needed in four areas to realize the potential of personalized medicine: (i) processing large-scale robust genomic data; (ii) interpreting the functional effect and the impact of genomic variation; (iii) integrating systems data to relate complex genetic interactions with phenotypes; and (iv) translating these discoveries into medical practice. CONTACT: [email protected] Guy Haskin Fernald, Emidio Capriotti, Roxana Daneshjou, Konrad J. Karczewski, Russ B. Altman |
Bioinform. | 5 |
| 2011 | CAMPAIGN: an open-source library of GPU-accelerated data clustering algorithmsabstractMOTIVATION: Data clustering techniques are an essential component of a good data analysis toolbox. Many current bioinformatics applications are inherently compute-intense and work with very large datasets. Sequential algorithms are inadequate for providing the necessary performance. For this reason, we have created Clustering Algorithms for Massively Parallel Architectures, Including GPU Nodes (CAMPAIGN), a central resource for data clustering algorithms and tools that are implemented specifically for execution on massively parallel processing architectures. RESULTS: CAMPAIGN is a library of data clustering algorithms and tools, written in 'C for CUDA' for Nvidia GPUs. The library provides up to two orders of magnitude speed-up over respective CPU-based clustering algorithms and is intended as an open-source resource. New modules from the community will be accepted into the library and the layout of it is such that it can easily be extended to promising future platforms such as OpenCL. AVAILABILITY: Releases of the CAMPAIGN library are freely available for download under the LGPL from https://simtk.org/home/campaign. Source code can also be obtained through anonymous subversion access as described on https://simtk.org/scm/?group_id=453. CONTACT: [email protected]. Kai Kohlhoff, Marc Sosnick-Pérez, William T. Hsu, Vijay S. Pande, Russ B. Altman |
Bioinform. | 5 |
| 2011 | Improving the prediction of disease-related variants using protein three-dimensional structureabstractBACKGROUND: Single Nucleotide Polymorphisms (SNPs) are an important source of human genome variability. Non-synonymous SNPs occurring in coding regions result in single amino acid polymorphisms (SAPs) that may affect protein function and lead to pathology. Several methods attempt to estimate the impact of SAPs using different sources of information. Although sequence-based predictors have shown good performance, the quality of these predictions can be further improved by introducing new features derived from three-dimensional protein structures. RESULTS: In this paper, we present a structure-based machine learning approach for predicting disease-related SAPs. We have trained a Support Vector Machine (SVM) on a set of 3,342 disease-related mutations and 1,644 neutral polymorphisms from 784 protein chains. We use SVM input features derived from the protein's sequence, structure, and function. After dataset balancing, the structure-based method (SVM-3D) reaches an overall accuracy of 85%, a correlation coefficient of 0.70, and an area under the receiving operating characteristic curve (AUC) of 0.92. When compared with a similar sequence-based predictor, SVM-3D results in an increase of the overall accuracy and AUC by 3%, and correlation coefficient by 0.06. The robustness of this improvement has been tested on different datasets and in all the cases SVM-3D performs better than previously developed methods even when compared with PolyPhen2, which explicitly considers in input protein structure information. CONCLUSION: This work demonstrates that structural information can increase the accuracy of disease-related SAPs identification. Our results also quantify the magnitude of improvement on a large dataset. This improvement is in agreement with previously observed results, where structure information enhanced the prediction of protein stability changes upon mutation. Although the structural information contained in the Protein Data Bank is limiting the application and the performance of our structure-based method, we expect that SVM-3D will result in higher accuracy when more structural date become available. Emidio Capriotti, Russ B. Altman |
BMC Bioinform. | 2 |
| 2011 | 2010 Translational bioinformatics year in reviewabstractA review of 2010 research in translational bioinformatics provides much to marvel at. We have seen notable advances in personal genomics, pharmacogenetics, and sequencing. At the same time, the infrastructure for the field has burgeoned. While acknowledging that, according to researchers, the members of this field tend to be overly optimistic, the authors predict a bright future. Russ B. Altman, Katharine S. Miller |
J. Am. Medical Informatics Assoc. | 1 |
| 2011 | Using Multiple Microenvironments to Find Similar Ligand-Binding Sites: Application to Kinase Inhibitor BindingabstractThe recognition of cryptic small-molecular binding sites in protein structures is important for understanding off-target side effects and for recognizing potential new indications for existing drugs. Current methods focus on the geometry and detailed chemical interactions within putative binding pockets, but may not recognize distant similarities where dynamics or modified interactions allow one ligand to bind apparently divergent binding pockets. In this paper, we introduce an algorithm that seeks similar microenvironments within two binding sites, and assesses overall binding site similarity by the presence of multiple shared microenvironments. The method has relatively weak geometric requirements (to allow for conformational change or dynamics in both the ligand and the pocket) and uses multiple biophysical and biochemical measures to characterize the microenvironments (to allow for diverse modes of ligand binding). We term the algorithm PocketFEATURE, since it focuses on pockets using the FEATURE system for characterizing microenvironments. We validate PocketFEATURE first by showing that it can better discriminate sites that bind similar ligands from those that do not, and by showing that we can recognize FAD-binding sites on a proteome scale with Area Under the Curve (AUC) of 92%. We then apply PocketFEATURE to evolutionarily distant kinases, for which the method recognizes several proven distant relationships, and predicts unexpected shared ligand binding. Using experimental data from ChEMBL and Ambit, we show that at high significance level, 40 kinase pairs are predicted to share ligands. Some of these pairs offer new opportunities for inhibiting two proteins in a single pathway. Tianyun Liu, Russ B. Altman |
PLoS Comput. Biol. | 2 |
| 2011 | Fast Flexible Modeling of RNA Structure Using Internal CoordinatesabstractModeling the structure and dynamics of large macromolecules remains a critical challenge. Molecular dynamics (MD) simulations are expensive because they model every atom independently, and are difficult to combine with experimentally derived knowledge. Assembly of molecules using fragments from libraries relies on the database of known structures and thus may not work for novel motifs. Coarse-grained modeling methods have yielded good results on large molecules but can suffer from difficulties in creating more detailed full atomic realizations. There is therefore a need for molecular modeling algorithms that remain chemically accurate and economical for large molecules, do not rely on fragment libraries, and can incorporate experimental information. RNABuilder works in the internal coordinate space of dihedral angles and thus has time requirements proportional to the number of moving parts rather than the number of atoms. It provides accurate physics-based response to applied forces, but also allows user-specified forces for incorporating experimental information. A particular strength of RNABuilder is that all Leontis-Westhof basepairs can be specified as primitives by the user to be satisfied during model construction. We apply RNABuilder to predict the structure of an RNA molecule with 160 bases from its secondary structure, as well as experimental information. Our model matches the known structure to 10.2 Angstroms RMSD and has low computational expense. Samuel Flores, Michael A. Sherman, Christopher M. Bruns, Peter K. Eastman, Russ B. Altman |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2010 | Editorial: Current progress in Bioinformatics 2010abstractIn this issue, we provide the next installment of our annual feature ‘Current Progress in Bioinformatics’. Each year, we invite leaders in bioinformatics to summarize the progress and challenges in hot and emerging areas of inquiry. We ask the authors specifically to summarize the past 12–18 months of important literature, and to provide a perspective in the field. The fields are chosen based on many different considerations: we start by studying the newest work in both bioinformatics journals as well as application-area journals in biology and medicine. We also consider the papers presented at bioinformatics conferences throughout the year, and the general ‘buzz’ of scientific areas. This year, the areas really jumped out without us having to work too hard. The field of bioinformatics is now a relatively stable discipline, with journals, conferences and a professional society all going back more than 10 years. Nonetheless, it is amazing to see how certain problems can ‘sneak up’ on us—the emergence of the data handling and analysis challenges associated with next-generation sequencing data is a prime example—many of us would have said that sequence analysis and genome assembly was a relatively dormant field, having solved most of the key problems for analysis and search that confronted biologists. Wow, were we wrong! Within months, many biological lab infrastructures were brought to their knees by the volume of data created by the new sequencing machines. Basic capabilities like storage, raw data triage and certainly search all were stressed by these data and the issues of basic sequence analysis have bounced back as a very busy area of bioinformatics. Similar ‘surprises’ have emerged from translational bioinformatics, synthetic biology and others. So we are pleased to have selected a few key areas where we can present a summary of progress. As usual, we organize these reviews in the rough order of the central dogma from DNA to RNA to proteins to networks to organs to organisms. Dalca and Brudno present a summary of next-generation sequencing informatics challenges in ‘Genome variation discovery with next generation sequencing data’. The throughput of these new sequencing technologies is breathtaking, but so are the challenges in terms of noise. After first providing a brief review of the basic sequencing technologies, they review methods for mapping reads on to existing assemblies, SNP and indel discovery, and discovery of structural variation on a larger scale. They also mention software systems that are emerging for providing a platform for the continuing analysis of these important data. The study of biology from a systems perspective has emerged as an important theme of bioinformatics in the last 5 years (often under the rubric of ‘systems biology’ but more generally as network and graph-based views of biological systems). The initial approaches often used static views of interactions in an ‘equilibrium’ view of biological homeostasis. In ‘Towards the dynamic interactome: it’s about time’ Przytycka, Singh and Slonim provide a review of recent methods to look at the evolution of interactions over time. Clearly, in the context of developmental biology (including stem cell biology) and dynamic response to perturbations such as drugs or other environmental inputs, it is absolutely critical to understand the dynamic response modes of cellular networks. In this review, the authors review the importance of time, space and context in understanding biological responses, the data sources that are relevant to these questions and the informatics techniques that are useful. It is exciting to imagine our static pathways ‘coming to life’ as we overlay data sets that shed light on dynamics. They end with the suggestion that our understanding of human genetic variation will necessarily require dynamic models of human biology. A critical technical capability for the analysis of biological systems has become the ability to integrate large data sets. The initial enthusiasm for any particular type of data (e.g. expression data) typically yields to the recognition that biological inferences are most robust when they stem from multiple data sources with different sources of bias and noise. In ‘Knowledge-based data analysis comes of age’, Michael Ochs summarizes the progress in overcoming the curse of dimensionality by using powerful mathematical (often explicitly probabilistic, Bayesian) models to normatively combine our a priori probabilities with data to extract accurate a posteriori estimates of biological parameters of interest. The high false positive rate of many high-throughput experimental technologies demands robust formal methods for integration of data. In ‘Pathway tools version 13.0: integrated software for pathway/genome informatics and systems biology’, Karp and coworkers provide a useful update on the latest updates to the Pathway Tools environment for capturing and analyzing genetic, protein, metabolic and regulatory networks in over 800 organisms. This provides a key integrating piece of technology for the study of molecular systems biology. The ultimate test of our understanding of biological systems is to physically manipulate them to create new capabilities. Many consider the emerging field of ‘synthetic biology’ to be an entirely experimental venture, but Alterovitz, Muso and Ramoni in ‘The challenges of informatics in synthetic biology: from biomolecular networks to artificial organisms’ argue that in silico support for synthetic biology is absolutely critical and forms the beginning of the CAD/CAM (computer-aided design/computer-aided manufacturing) era of biology and bioinformatics! To get things started, the authors review the emergence of standards for describing biological parts, and the major experimental areas of investigation, including synthesis of DNA sequences, gene expression control constructs and systems biology approaches to engineering new pathways. The rise in interest in ‘translational research’ in which basic science insights are brought to impact patient care has naturally lead to the emergence of ‘translational bioinformatics’ in which bioinformatics tools are used to study clinical/medical phenomena. Translational bioinformatics can be used to characterize the genetic underpinnings of disease, as well as to understand the molecular response to drugs, and (more generally) the link between molecular biology and clinical phenotypes. In ‘Advances in translational bioinformatics: computational approaches for the hunting of disease genes’, Kann provides an overview of how the combination of genome sequencing and high-throughput functional genomics data has enabled new understanding of the molecular bases for disease and its therapy. She summarizes how bioinformatics technologies can find new genetic modulators of disease, as well as the important variations associated with these genes. In many ways, one ultimate goal of bioinformatics is a working in silico model of a cell, organ or organism so that medical interventions can be tested on models before they are used on humans. The field of patient-specific modeling is just now emerging, and Neal and Kerckhoffs provide a summary in ‘Current progress in patient-specific modeling’. Currently, the field is dominated by ‘top down’ models based on 3D image data. Important progress has been made in modeling blood vessels and cardiovascular physiology, the heart (particularly for device modeling) and the musculoskeletal system. The opportunities for creating multiscale models that link molecular information to these higher level physiological systems are myriad. Not part of our ‘current progress’ series, but also included in this issue are four valuable additional reviews. Duval and Hao discuss recent progress in gene selection for accurate classification. Sloot and Hoekstra review multiscale modeling in a piece that is very complementary to the one by Neal and Kerckhoffs. The review of cellular reaction systems by Kirkilionis beautifully complements that of Przytycka, Singh and Slonim. Finally, Kim and Park provide a useful review of a new book ‘Modern Genome Annotation’. In summary, the field of bioinformatics continues to evolve to face the challenges that emerge across biology and medicine. We have shown remarkably flexibility in focusing on the data analysis needs of these communities, even when they take us a little by surprise! We hope you will enjoy reading about these important trends, and imagining where they will take us in the coming years. Russ B. Altman |
Briefings Bioinform. | 1 |
| 2010 | Content-based microarray search using differential expression profilesabstractBACKGROUND: With the expansion of public repositories such as the Gene Expression Omnibus (GEO), we are rapidly cataloging cellular transcriptional responses to diverse experimental conditions. Methods that query these repositories based on gene expression content, rather than textual annotations, may enable more effective experiment retrieval as well as the discovery of novel associations between drugs, diseases, and other perturbations. RESULTS: We develop methods to retrieve gene expression experiments that differentially express the same transcriptional programs as a query experiment. Avoiding thresholds, we generate differential expression profiles that include a score for each gene measured in an experiment. We use existing and novel dimension reduction and correlation measures to rank relevant experiments in an entirely data-driven manner, allowing emergent features of the data to drive the results. A combination of matrix decomposition and p-weighted Pearson correlation proves the most suitable for comparing differential expression profiles. We apply this method to index all GEO DataSets, and demonstrate the utility of our approach by identifying pathways and conditions relevant to transcription factors Nanog and FoxO3. CONCLUSIONS: Content-based gene expression search generates relevant hypotheses for biological inquiry. Experiments across platforms, tissue types, and protocols inform the analysis of new datasets. Jesse M. Engreitz, Alexander A. Morgan, Joel Dudley, Rong Chen 0006, Rahul Thathoo, Russ B. Altman, Atul J. Butte |
BMC Bioinform. | 6 |
| 2010 | An integrative method for scoring candidate genes from association studies: application to warfarin dosingabstractBACKGROUND: A key challenge in pharmacogenomics is the identification of genes whose variants contribute to drug response phenotypes, which can include severe adverse effects. Pharmacogenomics GWAS attempt to elucidate genotypes predictive of drug response. However, the size of these studies has severely limited their power and potential application. We propose a novel knowledge integration and SNP aggregation approach for identifying genes impacting drug response. Our SNP aggregation method characterizes the degree to which uncommon alleles of a gene are associated with drug response. We first use pre-existing knowledge sources to rank pharmacogenes by their likelihood to affect drug response. We then define a summary score for each gene based on allele frequencies and train linear and logistic regression classifiers to predict drug response phenotypes. RESULTS: We applied our method to a published warfarin GWAS data set comprising 181 individuals. We find that our method can increase the power of the GWAS to identify both VKORC1 and CYP2C9 as warfarin pharmacogenes, where the original analysis had only identified VKORC1. Additionally, we find that our method can be used to discriminate between low-dose (AUROC=0.886) and high-dose (AUROC=0.764) responders. CONCLUSIONS: Our method offers a new route for candidate pharmacogene discovery from pharmacogenomics GWAS, and serves as a foundation for future work in methods for predictive pharmacogenomics. Nicholas P. Tatonetti, Joel Dudley, Hersh Sagreiya, Atul J. Butte, Russ B. Altman |
BMC Bioinform. | 5 |
| 2010 | Using text to build semantic networks for pharmacogenomics
Adrien Coulet, Nigam H. Shah, Yael Garten, Mark A. Musen, Russ B. Altman |
J. Biomed. Informatics | 5 |
| 2010 | Independent component analysis: Mining microarray data for fundamental human gene expression modules
Jesse M. Engreitz, Bernie J. Daigle Jr., Jonathan J. Marshall, Russ B. Altman |
J. Biomed. Informatics | 4 |
| 2010 | The utility of general purpose versus specialty clinical databases for research: Warfarin dose estimation from extracted clinical variables
Hersh Sagreiya, Russ B. Altman |
J. Biomed. Informatics | 2 |
| 2010 | Using Pre-existing Microarray Datasets to Increase Experimental Power: Application to Insulin ResistanceabstractAlthough they have become a widely used experimental technique for identifying differentially expressed (DE) genes, DNA microarrays are notorious for generating noisy data. A common strategy for mitigating the effects of noise is to perform many experimental replicates. This approach is often costly and sometimes impossible given limited resources; thus, analytical methods are needed which increase accuracy at no additional cost. One inexpensive source of microarray replicates comes from prior work: to date, data from hundreds of thousands of microarray experiments are in the public domain. Although these data assay a wide range of conditions, they cannot be used directly to inform any particular experiment and are thus ignored by most DE gene methods. We present the SVD Augmented Gene expression Analysis Tool (SAGAT), a mathematically principled, data-driven approach for identifying DE genes. SAGAT increases the power of a microarray experiment by using observed coexpression relationships from publicly available microarray datasets to reduce uncertainty in individual genes' expression measurements. We tested the method on three well-replicated human microarray datasets and demonstrate that use of SAGAT increased effective sample sizes by as many as 2.72 arrays. We applied SAGAT to unpublished data from a microarray study investigating transcriptional responses to insulin resistance, resulting in a 50% increase in the number of significant genes detected. We evaluated 11 (58%) of these genes experimentally using qPCR, confirming the directions of expression change for all 11 and statistical significance for three. Use of SAGAT revealed coherent biological changes in three pathways: inflammation, differentiation, and fatty acid synthesis, furthering our molecular understanding of a type 2 diabetes risk factor. We envision SAGAT as a means to maximize the potential for biological discovery from subtle transcriptional responses, and we provide it as a freely available software package that is immediately applicable to any human microarray study. Bernie J. Daigle Jr., Alicia Deng, Tracey McLaughlin, Samuel W. Cushman, Margaret C. Cam, Gerald Reaven, Philip S. Tsao, Russ B. Altman |
PLoS Comput. Biol. | 8 |
| 2009 | A General Framework for Dose Optimization
Robert G. Turcott, Hersh Sagreiya, Euan A. Ashley, Russ B. Altman, Amar K. Das |
AMIA | 4 |
| 2009 | Knowledge-based instantiation of full atomic detail into coarse-grain RNA 3D structural modelsabstractMOTIVATION: The recent development of methods for modeling RNA 3D structures using coarse-grain approaches creates a need to bridge low- and high-resolution modeling methods. Although they contain topological information, coarse-grain models lack atomic detail, which limits their utility for some applications. RESULTS: We have developed a method for adding full atomic detail to coarse-grain models of RNA 3D structures. Our method [Coarse to Atomic (C2A)] uses geometries observed in known RNA crystal structures. Our method rebuilds full atomic detail from ideal coarse-grain backbones taken from crystal structures to within 1.87-3.31 A RMSD of the full atomic crystal structure. When starting from coarse-grain models generated by the modeling tool NAST, our method builds full atomic structures that are within 1.00 A RMSD of the starting structure. The resulting full atomic structures can be used as starting points for higher resolution modeling, thus bridging high- and low-resolution approaches to modeling RNA 3D structure. AVAILABILITY: Code for the C2A method, as well as the examples discussed in this article, are freely available at www.simtk.org/home/c2a. CONTACT: [email protected] Magdalena A. Jonikas, Randall J. Radmer, Russ B. Altman |
Bioinform. | 3 |
| 2009 | Pharmspresso: a text mining tool for extraction of pharmacogenomic concepts and relationships from full textabstractBACKGROUND: Pharmacogenomics studies the relationship between genetic variation and the variation in drug response phenotypes. The field is rapidly gaining importance: it promises drugs targeted to particular subpopulations based on genetic background. The pharmacogenomics literature has expanded rapidly, but is dispersed in many journals. It is challenging, therefore, to identify important associations between drugs and molecular entities--particularly genes and gene variants, and thus these critical connections are often lost. Text mining techniques can allow us to convert the free-style text to a computable, searchable format in which pharmacogenomic concepts (such as genes, drugs, polymorphisms, and diseases) are identified, and important links between these concepts are recorded. Availability of full text articles as input into text mining engines is key, as literature abstracts often do not contain sufficient information to identify these pharmacogenomic associations. RESULTS: Thus, building on a tool called Textpresso, we have created the Pharmspresso tool to assist in identifying important pharmacogenomic facts in full text articles. Pharmspresso parses text to find references to human genes, polymorphisms, drugs and diseases and their relationships. It presents these as a series of marked-up text fragments, in which key concepts are visually highlighted. To evaluate Pharmspresso, we used a gold standard of 45 human-curated articles. Pharmspresso identified 78%, 61%, and 74% of target gene, polymorphism, and drug concepts, respectively. CONCLUSION: Pharmspresso is a text analysis tool that extracts pharmacogenomic concepts from the literature automatically and thus captures our current understanding of gene-drug interactions in a computable form. We have made Pharmspresso available at http://pharmspresso.stanford.edu. Yael Garten, Russ B. Altman |
BMC Bioinform. | 2 |
| 2008 | M-BISON: Microarray-based integration of data sources using networksabstractBACKGROUND: The accurate detection of differentially expressed (DE) genes has become a central task in microarray analysis. Unfortunately, the noise level and experimental variability of microarrays can be limiting. While a number of existing methods partially overcome these limitations by incorporating biological knowledge in the form of gene groups, these methods sacrifice gene-level resolution. This loss of precision can be inappropriate, especially if the desired output is a ranked list of individual genes. To address this shortcoming, we developed M-BISON (Microarray-Based Integration of data SOurces using Networks), a formal probabilistic model that integrates background biological knowledge with microarray data to predict individual DE genes. RESULTS: M-BISON improves signal detection on a range of simulated data, particularly when using very noisy microarray data. We also applied the method to the task of predicting heat shock-related differentially expressed genes in S. cerevisiae, using an hsf1 mutant microarray dataset and conserved yeast DNA sequence motifs. Our results demonstrate that M-BISON improves the analysis quality and makes predictions that are easy to interpret in concert with incorporated knowledge. Specifically, M-BISON increases the AUC of DE gene prediction from .541 to .623 when compared to a method using only microarray data, and M-BISON outperforms a related method, GeneRank. Furthermore, by analyzing M-BISON predictions in the context of the background knowledge, we identified YHR124W as a potentially novel player in the yeast heat shock response. CONCLUSION: This work provides a solid foundation for the principled integration of imperfect biological knowledge with gene expression data and other high-throughput data sources. Bernie J. Daigle Jr., Russ B. Altman |
BMC Bioinform. | 2 |
| 2008 | MScanner: a classifier for retrieving Medline citationsabstractBACKGROUND: Keyword searching through PubMed and other systems is the standard means of retrieving information from Medline. However, ad-hoc retrieval systems do not meet all of the needs of databases that curate information from literature, or of text miners developing a corpus on a topic that has many terms indicative of relevance. Several databases have developed supervised learning methods that operate on a filtered subset of Medline, to classify Medline records so that fewer articles have to be manually reviewed for relevance. A few studies have considered generalisation of Medline classification to operate on the entire Medline database in a non-domain-specific manner, but existing applications lack speed, available implementations, or a means to measure performance in new domains. RESULTS: MScanner is an implementation of a Bayesian classifier that provides a simple web interface for submitting a corpus of relevant training examples in the form of PubMed IDs and returning results ranked by decreasing probability of relevance. For maximum speed it uses the Medical Subject Headings (MeSH) and journal of publication as a concise document representation, and takes roughly 90 seconds to return results against the 16 million records in Medline. The web interface provides interactive exploration of the results, and cross validated performance evaluation on the relevant input against a random subset of Medline. We describe the classifier implementation, cross validate it on three domain-specific topics, and compare its performance to that of an expert PubMed query for a complex topic. In cross validation on the three sample topics against 100,000 random articles, the classifier achieved excellent separation of relevant and irrelevant article score distributions, ROC areas between 0.97 and 0.99, and averaged precision between 0.69 and 0.92. CONCLUSION: MScanner is an effective non-domain-specific classifier that operates on the entire Medline database, and is suited to retrieving topics for which many features may indicate relevance. Its web interface simplifies the task of classifying Medline citations, compared to building a pre-filter and classifier specific to the topic. The data sets and open source code used to obtain the results in this paper are available on-line and as supplementary material, and the web interface may be accessed at http://mscanner.stanford.edu. Graham L. Poulter, Daniel L. Rubin, Russ B. Altman, Cathal Seoighe |
BMC Bioinform. | 3 |
| 2008 | The Simbios National Center: Systems Biology in MotionabstractPhysics-based simulation is needed to understand the function of biological structures and can be applied across a wide range of scales, from molecules to organisms. Simbios (the National Center for Physics-Based Simulation of Biological Structures, http://www.simbios.stanford.edu/) is one of seven NIH-supported National Centers for Biomedical Computation. This article provides an overview of the mission and achievements of Simbios, and describes its place within systems biology. Understanding the interactions between various parts of a biological system and integrating this information to understand how biological systems function is the goal of systems biology. Many important biological systems comprise complex structural systems whose components interact through the exchange of physical forces, and whose movement and function is dictated by those forces. In particular, systems that are made of multiple identifiable components that move relative to one another in a constrained manner are multibody systems. Simbios' focus is creating methods for their simulation. Simbios is also investigating the biomechanical forces that govern fluid flow through deformable vessels, a central problem in cardiovascular dynamics. In this application, the system is governed by the interplay of classical forces, but the motion is distributed smoothly through the materials and fluids, requiring the use of continuum methods. In addition to the research aims, Simbios is working to disseminate information, software and other resources relevant to biological systems in motion. Jeanette P. Schmidt, Scott L. Delp, Michael A. Sherman, Charles A. Taylor, Vijay S. Pande, Russ B. Altman |
Proc. IEEE | 6 |
| 2008 | Efficient Algorithms to Explore Conformation Spaces of Flexible Protein LoopsabstractSeveral applications in biology - e.g., incorporation of protein flexibility in ligand docking algorithms, interpretation of fuzzy X-ray crystallographic data, and homology modeling - require computing the internal parameters of a flexible fragment (usually, a loop) of a protein in order to connect its termini to the rest of the protein without causing any steric clash. One must often sample many such conformations in order to explore and adequately represent the conformational range of the studied loop. While sampling must be fast, it is made difficult by the fact that two conflicting constraints - kinematic closure and clash avoidance - must be satisfied concurrently. This paper describes two efficient and complementary sampling algorithms to explore the space of closed clash-free conformations of a flexible protein loop. The "seed sampling" algorithm samples broadly from this space, while the "deformation sampling" algorithm uses seed conformations as starting points to explore the conformation space around them at a finer grain. Computational results are presented for various loops ranging from 5 to 25 residues. More specific results also show that the combination of the sampling algorithms with a functional site prediction software (FEATURE) makes it possible to compute and recognize calcium-binding loop conformations. The sampling algorithms are implemented in a toolkit (LoopTK), which is available at https://simtk.org/home/looptk. Peggy Yao, Ankur Dhanik, Nathan Marz, Ryan Propper, Charles Kou, Guanfeng Liu 0001, Henry van den Bedem, Jean-Claude Latombe, Inbal Halperin-Landsberg, Russ B. Altman |
IEEE ACM Trans. Comput. Biol. Bioinform. | 10 |
| 2007 | Combining Simulation and Machine Learning to Recognize Function in 4DabstractThis paper is a talk by Russ Biagio Altman. It discusses structure-based protein function annotation using machine learning, physics-based simulation of structure, and how they can be profitably combined to improve our understanding of molecular structure and function. Russ B. Altman |
BIBM | 1 |
| 2007 | Current progress in bioinformatics 2007abstractBriefings in Bioinformatics is pleased to present our third annual ‘Current Progress in Bioinformatics’ special issue. As in previous years, we have attempted to identify exciting or emerging fields of bioinformatics, and have asked leaders in these fields to present a brief summary of progress over the last 18–24 months and an annotated biography drawing attention to papers of particular significance. Each year, we have a logistical task of setting the order of the articles to appear in this volume. Typically, we organize them based on the linear logic of biology's central dogma: from DNA to RNA to protein to function and phenotype. The central dogma has undergone a transformation in the last 10 years, however. Biologists have demonstrated that the simple, linear model must be augmented with multiple feedback loops. For example, RNA feeds back to affect gene regulation, and feeds forward to modulate protein function. RNA itself is modified by proteins that can alter the message. The linear central dogma has become the networked central dogma! Our field of bioinformatics is also starting to become a network. The relationship between different subdisciplines is getting increasingly complex, and promises to keep us all busy for some time. So what is the appropriate order of seven reviews on (i) metabolomics, (ii) structured RNA, (iii) proteomics, (iv) gene product networks, (v) proteins and disease, (vi) biodiversity and (vii) text mining? Well, we chose that order. Hopefully it makes sense, but probably it doesn't matter! In the first review, David Wishart provides a summary on ‘Current Progress in Computational Metabolomics.’ He first provides a useful introduction to metabolomics—the study of the small molecules that interact with biological macromolecules as one of the major ‘omics’ pillars. As expected, many of the challenges to metabolomics are analogous to similar challenges in other branches of bioinformatics—the need to catalog small molecules in databases, to search, compare and classify them. In addition, there are important vocabulary and standards issues. There are particular challenges relating to the experimental reality of proteomics, where analytic measurements require special purpose software for laboratory information management and interpretation. Metabolomics is a field where chemoinformatics touches genomics. Thus, we get a glimpse of a future where informatics tools from neighboring disciplines are interoperable and create a potent infrastructure for discovery and engineering. Alain Laederach next reports on ‘Informatics challenges in Structured RNA.’ Understanding the protean (!) functions of RNA is a new challenge for computational biology and bioinformatics. In particular, the field is approaching challenges associated with understanding the physical properties of 3D RNA molecules, which (like proteins) fold into precise 3D shapes, catalyze many important reactions, and participate in the control of gene expression. Unlike proteins, however, they are made of four subunits (not 20), are dominated by electrostatics (RNA itself is extremely electronegative), and form their secondary structure almost entirely before assembling into their final 3D structure. Special purpose methods, therefore, are required to analyze the experimental data about their folding kinetics, the nature of their thermodynamic stability, and to understand how to use RNA 3D structure to understand function when the 1D sequence alone is insufficient. Bobbie-Jo Webb-Robertson writes about ‘Current Trends in Computational Inference from Mass Spectrometry-based Proteomics.’ She focuses, in particular, on recent achievements in mass-spectrometry (MS) based proteomics. The power of MS for interrogating cellular protein populations is immense—it can be used to identify proteins, characterize new proteins by de novo sequencing, characterize posttranslationally modified proteins, quantify proteins in the cellular milieu and assess protein–protein interactions. As high-throughput biology extends from genome to transcriptome to proteome, the richness of information with direct relevance to phenotype is exciting. Of course, systems biological models should benefit greatly from proteomic technologies—which provides parameters for the models and is useful for validation. Balaji Srinivasan's contribution, ‘Toward Reference Networks for Key Model Organisms’ reviews the key data sources used to build interaction networks, and then discusses progress in the algorithms for comparing the resulting networks. In particular, he discusses methods for aligning networks to identify conserved network modules as well as those that confer species-specific capabilities. The challenges involve integration of multiple data sets, dealing with the differential noise in these data sets, creating robust data structures for representation and visualizing these networks. Interestingly, these activities highlight the need for a network ontology (NO!) that provides standard terms for annotating interaction networks. Maricel Kann provides an overview of work in ‘Protein Interactions and Disease: Computational Approaches to Uncover the Etiology of Diseases.’ She focuses on the role of proteins in disease, ranging from individual protein mutations, to protein ensembles, to networks of interacting proteins. In each of these there are difficult challenges, in part stemming from the hierarchical relationship of individual molecular mutations to emergent functional properties of the cell (and organism). Until recently, protein–interaction data was not amenable to high-throughput experimental measurements, but their emergence promises a connection between structural bioinformatics and systems biology. Indra Sarkar provides a fascinating introduction to a fascinating emerging field in ‘Biodiversity Informatics: Organizing and Linking Information Across the Spectrum of Life.’ Rising out of multiple intellectual threads, this currently focuses on methods for generating reliable species identifiers. With increasing interest in complete characterization of the species found in ecological niches (consider, for example, the metagenomic sequencing of entire bacterial ecosystems), nomenclatures for species have become critical. At the same time, we need to integrate hundreds of years of taxonomic and biological phenotypic descriptions into the digital corpus. The first critical achievements are the creation of systems for naming species, resolving alternative names and querying large heterogeneous data sources with them. Species nomenclatures are only the beginning, as we also need to characterize their niches—climates, geography, disease and other interacting species. The diversity of challenges is truly impressive. Finally, Pierre Zweigenbaum and colleagues provide a comprehensive review of recent biological natural language processing (BioNLP) in ‘New frontiers for biomedical text mining: current progress.’ Biological text analysis is an important activity in bioinformatics. Despite the obvious drawbacks, scientists continue to publish their scientific findings using natural language—with all its ambiguities and subtleties. Worse yet—scientists are creating and reporting useful knowledge at rates far beyond what can be read even by the most motivated. There has been outstanding progress in some of the ‘traditional’ areas of BioNLP—to the point where the authors suggest that some problems are solved or nearly solved. They introduce fascinating new areas, such as mining the figures in scientific papers, supporting the activities of database curators and mapping text to ontologies. Together, these seven reviews offer an exciting view of a field that continues to grow and extend its influence to all areas of biomedicine, creating a network of useful methods and data structures that will enable the next generation of discovery and engineering. Russ B. Altman |
Briefings Bioinform. | 1 |
| 2007 | Clustering protein environments for function prediction: finding PROSITE motifs in 3DabstractBACKGROUND: Structural genomics initiatives are producing increasing numbers of three-dimensional (3D) structures for which there is little functional information. Structure-based annotation of molecular function is therefore becoming critical. We previously presented FEATURE, a method for describing microenvironments around functional sites in proteins. However, FEATURE uses supervised machine learning and so is limited to building models for sites of known importance and location. We hypothesized that there are a large number of sites in proteins that are associated with function that have not yet been recognized. Toward that end, we have developed a method for clustering protein microenvironments in order to evaluate the potential for discovering novel sites that have not been previously identified. RESULTS: We have prototyped a computational method for rapid clustering of millions of microenvironments in order to discover residues whose surrounding environments are similar and which may therefore share a functional or structural role. We clustered nearly 2,000,000 environments from 9,600 protein chains and defined 4,550 clusters. As a preliminary validation, we asked whether known 3D environments associated with PROSITE motifs were "rediscovered". We found examples of clusters highly enriched for residues that share PROSITE sequence motifs. CONCLUSION: Our results demonstrate that we can cluster protein environments successfully using a simplified representation and K-means clustering algorithm. The rediscovery of known 3D motifs allows us to calibrate the size and intercluster distances that characterize useful clusters. This information will then allow us to find new clusters with similar characteristics that represent novel structural or functional sites. Sungroh Yoon, Jessica C. Ebert, Eui-Young Chung, Giovanni De Micheli, Russ B. Altman |
BMC Bioinform. | 5 |
| 2007 | Biomedical informatics training at Stanford in the 21st century
Russ B. Altman, Teri E. Klein |
J. Biomed. Informatics | 1 |
| 2007 | The International Society for Computational Biology 10th AnniversaryabstractPLoS Computational Biology is the official journal of the International Society for Computational Biology (ISCB), a partnership that was formed during the Journal's conception in 2005. With ISCB being the only international body representing computational biologists, it made perfect sense for PLoS Computational Biology to be closely affiliated. The Society had to take more of a chance than similar societies, choosing to step away from an existing financially beneficial subscription journal to align with an open access publication as a matter of principle. To our knowledge, ISCB was the first major international scientific society to do so.
Now, as PLoS Computational Biology reaches its two-year mark, ISCB simultaneously celebrates its tenth anniversary, having formed officially on June 18, 1997. We early presidents of ISCB reflect on the state of computational biology ten years ago, how far we have come since, and what thought-provoking future challenges might lie ahead with regard to innovations in publishing technologies. Lawrence Hunter, Russ B. Altman, Philip E. Bourne |
PLoS Comput. Biol. | 2 |
| 2006 | Annual Progress in Bioinformatics 2006abstractIn this issue, Briefings in Bioinformatics is happy to present the next installment of our special annual issue devoted to reviews of very active subdisciplines within our field. The editors surveyed recent publications in order to identify fields that are moving rapidly and would be good targets for summary and review. We asked authors to provide brief introductions to their field, and then to concentrate on contributions in the last 12–24 months of particular interest. In some cases, they also provided annotated bibliographies in which they highlighted papers of particularly high interest. The result is seven outstanding reviews. The influence of high-throughput genomic experimental techniques and the increasing interest on synthesis of information comes through strongly in this year's selections. We have ordered the reviews starting with those discussing tools close to the genome (the HapMap project, function prediction, graph methods for analyzing cellular networks), and then toward tool-oriented organization and sharing of knowledge (biological ontologies, the semantic web and open source software). The first set of reviews focuses on tools for understanding the central dogma and basic biological processes. The second set focuses on tools to assist scientists in the process of doing their work. In many ways, these are the two primary foci of bioinformatics, and it is reassuring to see that progress is balanced along both fronts. In the first review, Barnes provides an overview of the HapMap project for cataloging human genetic diversity. Understanding variation in the human genome is critical for understanding the variation in human phenotypes. The HapMap project is the natural follow-up to the human genome sequencing project, and seeks to characterize the variations in the human genome—initially in four groups of different geographic origin. The review describes the HapMap project's motivation, strategy, data resources and analytic challenges. Not surprisingly, variation in the human genome is not entirely independent, but shows a correlation structure (expressed as linkage disequilibrium or LD) that is critical for the design of studies that aim to understand the relationship between genotype and phenotype. In addition, this LD structure can be examined in the context of human population history to understand our origins. In the second review, Iddo Friedberg presents the challenges associated with annotating genes with their biological functions. It seems that nothing is easy about this task. First, genes are typically polyfunctional and therefore multiple experimental and theoretical sources are used to characterize their function. Annotation techniques must be careful to consider multiple sources of data, and must allow multiple annotations. Second, it is not clear what language should be used to describe gene function. Gene function may be understood in many contexts, and controlled terminologies are required in order to guarantee precise semantics when functional labels are used. Finally, the promise of computer algorithms for predicting function must be associated with gold standard methods to validate these predictions. The first two problems make this last one even more challenging. Aittokallio and Schwikowski present an overview of graph methods for biological networks in the third review. The availability of multiple high-throughput data sources using gene expression, proteomics, literature mining and other techniques provide information sufficient to create networks of interaction. These networks provide a global view of biological systems, and are the first step towards an integrated understanding of the emergent properties of these systems. Of course, the networks are often represented as graph, and require informatics tools for their analysis—often taking advantage of a mature computer science literature on graphical methods. The analyses that result are truly ‘multiscale’ because they range from global properties of the networks all the way to the analysis of individual interactions. A particular challenge is the identification of modules, clusters and recurring motifs that perform identifiable functions, and may be conserved across evolution. The authors also point out that integrated analyses that include multiple sources of data may provide better performance. Nearly every area of biomedical research is currently concerned about the integration, aggregation and annotation of experimental data and the associated knowledge. This concern stems from two observations: (i) the volume of data in most fields is exploding and is impossible to track manually and (ii) this data is useless if they cannot be indexed and retrieved with labels that are standardized and have clear semantics. Thus, biomedical ontologies have emerged as the primary hope for providing the required informatics infrastructure. In the fourth review, Bodenreider and Stevens describe the history of ontologies in biomedicine, and describe some remarkable developments in the last few years that have accelerated progress. Chief among these is the near-universal agreement that ontologies are a critical strategic need (even for individuals who had never heard of ontologies a decade ago), and that scientific institutions have been formed around this goal, introducing more resources and more constraints on their development. The progress on ontology is a critical prerequisite for the long-awaited emergence of the ‘semantic web for life science.’ The impact of the internet and world-wide-web on biomedical research has been profound, but some believe it could be even greater with better integration of tools, data and other resources—based on an ability to represent the semantics of these resources and link them appropriately. In the fifth review, Good and Wilkinson complement Bodenreider and Stevens by providing a ‘systems level’ view of progress towards building the next generation of web-based tools for scientists. An early success has been the rapid penetration of web services in bioinformatics, a low-cost method for making computational services available. The authors point out, however, that further progress may require a move from centralized control of data and algorithms toward a more open and inter-linked model. They argue that the primary challenges may be sociological and not technical. In the final sixth review, Stajich provides an overview of open-source software in bioinformatics. The success of the Linux operating system introduced the paradigm of open-source development, and many bioinformatics software developers embraced this paradigm as a way to accelerate progress in the field, by avoiding redundancy and promoting transparency. But has open-source software made an impact in bioinformatics, and how should individual developers decide whether to join an open-source project or build their own? This review provides several useful case studies to show the diversity of approaches that have lead to valuable software libraries covering a range of applications in molecular biology. The review also revisits a theme it shares with the articles on ontologies and the semantic web—the need for standards. Taken together, these last three reviews provide a fascinating view on how a maturing field of bioinformatics is organizing itself for the coming decades. In addition to the annual progress reviews, this issue includes an outstanding review of statistical methods for association studies by Montana. Increasingly, bioinformatics professionals are being asked to participate in projects with goals of associating genotype with phenotype. However, the literature on the analysis of genotype is rich and mature, and the complexities often overwhelm scientists who are otherwise very familiar with computing with DNA sequences. This review is a very useful primer on the state and progress of methods for statistical association analysis. We hope that you will find the six annual progress reviews and the review on statistical methods for genetic studies useful. Taken together, they demonstrate that bioinformatics continues to be an active and evolving field, responding to scientific and technical opportunities. Russ B. Altman |
Briefings Bioinform. | 1 |
| 2006 | Correction: Time to Organize the Bioinformatics Resourceome
Nicola Cannata, Emanuela Merelli, Russ B. Altman |
PLoS Comput. Biol. | 3 |
| 2005 | Editorial: Annual progress in bioinformatics
Russ B. Altman |
Briefings Bioinform. | 1 |
| 2005 | Research Paper: Using Petri Net Tools to Study Properties and Dynamics of Biological SystemsabstractPetri Nets (PNs) and their extensions are promising methods for modeling and simulating biological systems. We surveyed PN formalisms and tools and compared them based on their mathematical capabilities as well as by their appropriateness to represent typical biological processes. We measured the ability of these tools to model specific features of biological systems and answer a set of biological questions that we defined. We found that different tools are required to provide all capabilities that we assessed. We created software to translate a generic PN model into most of the formalisms and tools discussed. We have also made available three models and suggest that a library of such models would catalyze progress in qualitative modeling via PNs. Development and wide adoption of common formats would enable researchers to share models and use different tools to analyze them without the need to convert to proprietary formats. Mor Peleg, Daniel L. Rubin, Russ B. Altman |
J. Am. Medical Informatics Assoc. | 3 |
| 2005 | Application of Information Technology: A Statistical Approach to Scanning the Biomedical Literature for Pharmacogenetics KnowledgeabstractOBJECTIVE: Biomedical databases summarize current scientific knowledge, but they generally require years of laborious curation effort to build, focusing on identifying pertinent literature and data in the voluminous biomedical literature. It is difficult to manually extract useful information embedded in the large volumes of literature, and automated intelligent text analysis tools are becoming increasingly essential to assist in these curation activities. The goal of the authors was to develop an automated method to identify articles in Medline citations that contain pharmacogenetics data pertaining to gene-drug relationships. DESIGN: The authors built and evaluated several candidate statistical models that characterize pharmacogenetics articles in terms of word usage and the profile of Medical Subject Headings (MeSH) used in those articles. The best-performing model was used to scan the entire Medline article database (11 million articles) to identify candidate pharmacogenetics articles. RESULTS: A sampling of the articles identified from scanning Medline was reviewed by a pharmacologist to assess the precision of the method. The authors' approach identified 4,892 pharmacogenetics articles in the literature with 92% precision. Their automated method took a fraction of the time to acquire these articles compared with the time expected to be taken to accumulate them manually. The authors have built a Web resource (http://pharmdemo.stanford.edu/pharmdb/main.spy) to provide access to their results. CONCLUSION: A statistical classification approach can screen the primary literature to pharmacogenetics articles with high precision. Such methods may assist curators in acquiring pertinent literature in building biomedical databases. Daniel L. Rubin, Caroline F. Thorn, Teri E. Klein, Russ B. Altman |
J. Am. Medical Informatics Assoc. | 4 |
| 2005 | Time to Organize the Bioinformatics ResourceomeabstractThe initial steps toward a bioinformatics resourceome are clear. First, an overall ontology with the high-level concepts (algorithms, databases, organizations, papers, people, etc.) must be created, with a set of standard attributes and a standard set of relations between these concepts (e.g., people publish papers, papers describe algorithms or databases, organizations house people, etc.). The initial ontology should be compact and built for distributed collaborative extension. Second, a mechanism for people to extend this ontology with subconcepts in order to describe their own resources should be designed. The precise location of a tool within a taxonomy is not critical—the author will place it somewhere based on the location of similar/competing resources or based on a best-informed guess. Others may create links to the resource from other appropriate locations in the taxonomy in order to ensure that competing interpretations of the appropriate conceptual location for the resource are accommodated. Third, the formats for the ontologies and the resource descriptions should be published so enterprising software engineers can create interfaces for surfing, searching, and viewing the resources. The resulting distributed system of resource descriptions would be extensible, robust, and useful to the entire biomedical research community. Nicola Cannata, Emanuela Merelli, Russ B. Altman |
PLoS Comput. Biol. | 3 |
| 2004 | Editorial: Building successful biological databases
Russ B. Altman |
Briefings Bioinform. | 1 |
| 2004 | GAPSCORE: finding gene and protein names one word at a timeabstractMOTIVATION: New high-throughput technologies have accelerated the accumulation of knowledge about genes and proteins. However, much knowledge is still stored as written natural language text. Therefore, we have developed a new method, GAPSCORE, to identify gene and protein names in text. GAPSCORE scores words based on a statistical model of gene names that quantifies their appearance, morphology and context. RESULTS: We evaluated GAPSCORE against the Yapex data set and achieved an F-score of 82.5% (83.3% recall, 81.5% precision) for partial matches and 57.6% (58.5% recall, 56.7% precision) for exact matches. Since the method is statistical, users can choose score cutoffs that adjust the performance according to their needs. AVAILABILITY: GAPSCORE is available at http://bionlp.stanford.edu/gapscore/ Jeffrey T. Chang, Hinrich Schütze, Russ B. Altman |
Bioinform. | 3 |
| 2004 | Tools for loading MEDLINE into a local relational databaseabstractBACKGROUND: Researchers who use MEDLINE for text mining, information extraction, or natural language processing may benefit from having a copy of MEDLINE that they can manage locally. The National Library of Medicine (NLM) distributes MEDLINE in eXtensible Markup Language (XML)-formatted text files, but it is difficult to query MEDLINE in that format. We have developed software tools to parse the MEDLINE data files and load their contents into a relational database. Although the task is conceptually straightforward, the size and scope of MEDLINE make the task nontrivial. Given the increasing importance of text analysis in biology and medicine, we believe a local installation of MEDLINE will provide helpful computing infrastructure for researchers. RESULTS: We developed three software packages that parse and load MEDLINE, and ran each package to install separate instances of the MEDLINE database. For each installation, we collected data on loading time and disk-space utilization to provide examples of the process in different settings. Settings differed in terms of commercial database-management system (IBM DB2 or Oracle 9i), processor (Intel or Sun), programming language of installation software (Java or Perl), and methods employed in different versions of the software. The loading times for the three installations were 76 hours, 196 hours, and 132 hours, and disk-space utilization was 46.3 GB, 37.7 GB, and 31.6 GB, respectively. Loading times varied due to a variety of differences among the systems. Loading time also depended on whether data were written to intermediate files or not, and on whether input files were processed in sequence or in parallel. Disk-space utilization depended on the number of MEDLINE files processed, amount of indexing, and whether abstracts were stored as character large objects or truncated. CONCLUSIONS: Relational database (RDBMS) technology supports indexing and querying of very large datasets, and can accommodate a locally stored version of MEDLINE. RDBMS systems support a wide range of queries and facilitate certain tasks that are not directly supported by the application programming interface to PubMed. Because there is variation in hardware, software, and network infrastructures across sites, we cannot predict the exact time required for a user to load MEDLINE, but our results suggest that performance of the software is reasonable. Our database schemas and conversion software are publicly available at http://biotext.berkeley.edu. Diane E. Oliver, Gaurav Bhalotia, Ariel S. Schwartz, Russ B. Altman, Marti A. Hearst |
BMC Bioinform. | 4 |
| 2004 | White Paper: Training the Next Generation of Informaticians: The Impact of "BISTI" and Bioinformatics - A Report from the American College of Medical InformaticsabstractIn 2002-2003, the American College of Medical Informatics (ACMI) undertook a study of the future of informatics training. This project capitalized on the rapidly expanding interest in the role of computation in basic biological research, well characterized in the National Institutes of Health (NIH) Biomedical Information Science and Technology Initiative (BISTI) report. The defining activity of the project was the three-day 2002 Annual Symposium of the College. A committee, comprised of the authors of this report, subsequently carried out activities, including interviews with a broader informatics and biological sciences constituency, collation and categorization of observations, and generation of recommendations. The committee viewed biomedical informatics as an interdisciplinary field, combining basic informational and computational sciences with application domains, including health care, biological research, and education. Consequently, effective training in informatics, viewed from a national perspective, should encompass four key elements: (1). curricula that integrate experiences in the computational sciences and application domains rather than just concatenating them; (2). diversity among trainees, with individualized, interdisciplinary cross-training allowing each trainee to develop key competencies that he or she does not initially possess; (3). direct immersion in research and development activities; and (4). exposure across the wide range of basic informational and computational sciences. Informatics training programs that implement these features, irrespective of their funding sources, will meet and exceed the challenges raised by the BISTI report, and optimally prepare their trainees for careers in a field that continues to evolve. Charles P. Friedman, Russ B. Altman, Isaac S. Kohane, Kathleen A. McCormick, Perry L. Miller, Judy G. Ozbolt, Edward H. Shortliffe, Gary D. Stormo, M. Cleat Szczepaniak, David Tuck, Jeffrey J. Williamson |
J. Am. Medical Informatics Assoc. | 2 |
| 2003 | MutDB: annotating human variation with functionally relevant dataabstractSUMMARY: We have developed a resource, MutDB (http://mutdb.org/), to aid in determining which single nucleotide polymorphisms (SNPs) are likely to alter the function of their associated protein product. MutDB contains protein structure annotations and comparative genomic annotations for 8000 disease-associated mutations and SNPs found in the UCSC Annotated Genome and the human RefSeq gene set. MutDB provides interactive mutation maps at the gene and protein levels, and allows for ranking of their predicted functional consequences based on conservation in multiple sequence alignments. AVAILABILITY: http://mutdb.org/ SUPPLEMENTARY INFORMATION: http://mutdb.org/about/about.html Sean D. Mooney, Russ B. Altman |
Bioinform. | 2 |
| 2003 | A literature-based method for assessing the functional coherence of a gene groupabstractMOTIVATION: Many experimental and algorithmic approaches in biology generate groups of genes that need to be examined for related functional properties. For example, gene expression profiles are frequently organized into clusters of genes that may share functional properties. We evaluate a method, neighbor divergence per gene (NDPG), that uses scientific literature to assess whether a group of genes are functionally related. The method requires only a corpus of documents and an index connecting the documents to genes. RESULTS: We evaluate NDPG on 2796 functional groups generated by the Gene Ontology consortium in four organisms: mouse, fly, worm and yeast. NDPG finds functional coherence in 96, 92, 82 and 45% of the groups (at 99.9% specificity) in yeast, mouse, fly and worm respectively. Soumya Raychaudhuri, Russ B. Altman |
Bioinform. | 2 |
| 2003 | Knowledge acquisition, consistency checking and concurrency control for Gene Ontology (GO)abstractMOTIVATION: A critical element of the computational infrastructure required for functional genomics is a shared language for communicating biological data and knowledge. The Gene Ontology (GO; http://www.geneontology.org) provides a taxonomy of concepts and their attributes for annotating gene products. As GO increases in size, its ongoing construction and maintenance becomes more challenging. In this paper, we assess the applicability of a Knowledge Base Management System (KBMS), Protégé-2000, to the maintenance and development of GO. RESULTS: We transferred GO to Protégé-2000 in order to evaluate its suitability for GO. The graphical user interface supported browsing and editing of GO. Tools for consistency checking identified minor inconsistencies in GO and opportunities to reduce redundancy in its representation. The Protégé Axiom Language proved useful for checking ontological consistency. The PROMPT tool allowed us to track changes to GO. Using Protégé-2000, we tested our ability to make changes and extensions to GO to refine the semantics of attributes and classify more concepts. AVAILABILITY: Gene Ontology in Protégé-2000 and the associated code are located at http://smi.stanford.edu/projects/helix/gokbms/. Protégé-2000 is available from http://protege.stanford.edu. Iwei Yeh, Peter D. Karp, Natasha F. Noy, Russ B. Altman |
Bioinform. | 4 |
| 2003 | Inclusion of Textual Documentation in the Analysis of Multidimensional Data Sets: Application to Gene Expression Data
Soumya Raychaudhuri, Hinrich Schütze, Russ B. Altman |
Mach. Learn. | 3 |
| 2002 | Using binning to maintain confidentiality of medical data
Micheal Hewett, Russ B. Altman |
AMIA | 3 |
| 2002 | Representing genetic sequence data for pharmacogenomics: an evolutionary approach using ontological and relational modelsabstractAbstract Motivation: The information model chosen to store biological data affects the types of queries possible, database performance, and difficulty in updating that information model. Genetic sequence data for pharmacogenetics studies can be complex, and the best information model to use may change over time. As experimental and analytical methods change, and as biological knowledge advances, the data storage requirements and types of queries needed may also change. Results: We developed a model for genetic sequence and polymorphism data, and used XML Schema to specify the elements and attributes required for this model. We implemented this model as an ontology in a frame-based representation and as a relational model in a database system. We collected genetic data from two pharmacogenetics resequencing studies, and formulated queries useful for analysing these data. We compared the ontology and relational models in terms of query complexity, performance, and difficulty in changing the information model. Our results demonstrate benefits of evolving the schema for storing pharmacogenetics data: ontologies perform well in early design stages as the information model changes rapidly and simplify query formulation, while relational models offer improved query speed once the information model and types of queries needed stabilize. Availability: Our ontology and relational models are available at http://smi-web.stanford.edu/projects/helix/pubs/ismb02/. Contact: [email protected]@[email protected] Keywords: ontologies; relational databases; schema; data models; pharmacogenomics. Daniel L. Rubin, Farhad Shafa, Diane E. Oliver, Micheal Hewett, Russ B. Altman |
ISMB | 5 |
| 2002 | Modelling biological processes using workflow and Petri Net modelsabstractMOTIVATION: Biological processes can be considered at many levels of detail, ranging from atomic mechanism to general processes such as cell division, cell adhesion or cell invasion. The experimental study of protein function and gene regulation typically provides information at many levels. The representation of hierarchical process knowledge in biology is therefore a major challenge for bioinformatics. To represent high-level processes in the context of their component functions, we have developed a graphical knowledge model for biological processes that supports methods for qualitative reasoning. RESULTS: We assessed eleven diverse models that were developed in the fields of software engineering, business, and biology, to evaluate their suitability for representing and simulating biological processes. Based on this assessment, we combined the best aspects of two models: Workflow/Petri Net and a biological concept model. The Workflow model can represent nesting and ordering of processes, the structural components that participate in the processes, and the roles that they play. It also maps to Petri Nets, which allow verification of formal properties and qualitative simulation. The biological concept model, TAMBIS, provides a framework for describing biological entities that can be mapped to the workflow model. We tested our model by representing malaria parasites invading host erythrocytes, and composed queries, in five general classes, to discover relationships among processes and structural components. We used reachability analysis to answer queries about the dynamic aspects of the model. AVAILABILITY: The model is available at http://smi.stanford.edu/projects/helix/pubs/process-model/. Mor Peleg, Iwei Yeh, Russ B. Altman |
Bioinform. | 3 |
| 2002 | Nonparametric methods for identifying differentially expressed genes in microarray dataabstractMOTIVATION: Gene expression experiments provide a fast and systematic way to identify disease markers relevant to clinical care. In this study, we address the problem of robust identification of differentially expressed genes from microarray data. Differentially expressed genes, or discriminator genes, are genes with significantly different expression in two user-defined groups of microarray experiments. We compare three model-free approaches: (1). nonparametric t-test, (2). Wilcoxon (or Mann-Whitney) rank sum test, and (3). a heuristic method based on high Pearson correlation to a perfectly differentiating gene ('ideal discriminator method'). We systematically assess the performance of each method based on simulated and biological data under varying noise levels and p-value cutoffs. RESULTS: All methods exhibit very low false positive rates and identify a large fraction of the differentially expressed genes in simulated data sets with noise level similar to that of actual data. Overall, the rank sum test appears most conservative, which may be advantageous when the computationally identified genes need to be tested biologically. However, if a more inclusive list of markers is desired, a higher p-value cutoff or the nonparametric t-test may be appropriate. When applied to data from lung tumor and lymphoma data sets, the methods identify biologically relevant differentially expressed genes that allow clear separation of groups in question. Thus the methods described and evaluated here provide a convenient and robust way to identify differentially expressed genes for further biological and clinical analysis. Olga G. Troyanskaya, Mitchell E. Garber, Patrick O. Brown, David Botstein, Russ B. Altman |
Bioinform. | 5 |
| 2002 | Research Paper: Creating an Online Dictionary of Abbreviations from MEDLINEabstractOBJECTIVE: The growth of the biomedical literature presents special challenges for both human readers and automatic algorithms. One such challenge derives from the common and uncontrolled use of abbreviations in the literature. Each additional abbreviation increases the effective size of the vocabulary for a field. Therefore, to create an automatically generated and maintained lexicon of abbreviations, we have developed an algorithm to match abbreviations in text with their expansions. DESIGN: Our method uses a statistical learning algorithm, logistic regression, to score abbreviation expansions based on their resemblance to a training set of human-annotated abbreviations. We applied it to Medstract, a corpus of MEDLINE abstracts in which abbreviations and their expansions have been manually annotated. We then ran the algorithm on all abstracts in MEDLINE, creating a dictionary of biomedical abbreviations. To test the coverage of the database, we used an independently created list of abbreviations from the China Medical Tribune. MEASUREMENTS: We measured the recall and precision of the algorithm in identifying abbreviations from the Medstract corpus. We also measured the recall when searching for abbreviations from the China Medical Tribune against the database. RESULTS: On the Medstract corpus, our algorithm achieves up to 83% recall at 80% precision. Applying the algorithm to all of MEDLINE yielded a database of 781,632 high-scoring abbreviations. Of all the abbreviations in the list from the China Medical Tribune, 88% were in the database. CONCLUSION: We have developed an algorithm to identify abbreviations from text. We are making this available as a public abbreviation server at \url[http://abbreviation.stanford.edu/]. Jeffrey T. Chang, Hinrich Schütze, Russ B. Altman |
J. Am. Medical Informatics Assoc. | 3 |
| 2002 | Qualitative models of molecular function: linking genetic polymorphisms of tRNA to their functional sequelaeabstractThe exponential growth in the volume of biological information available makes it difficult for researchers to assemble the details into coherent models. Although an accurate model is ideal, full details are not generally available and are gained only incrementally. Therefore, as a first step toward integration of information, we propose a knowledge model for the qualitative representation of the relationships between mutations in genes and their effects at molecular cellular and clinical phenotypic levels. Our framework combines and extends two components: 1) a workflow model that allows hierarchical process and participant specifications; 2) Transparent Access to Multiple Bioinformatics Information Sources and the Unified Medical Language System, which serve as controlled biological and medical terminologies. By mapping our framework to Petri nets, we can perform qualitative simulations to validate models, and aid in predicting system behavior in the presence of dysfunctional components. This can be a step toward accurate quantitative models. Our application domain is the role of transfer ribonucleic acid molecules in protein translation-related disease. As an initial evaluation, we show that Petri nets derived from the historic and current views of the translation process yield different dynamic behavior. Our model is available at http://smi.stanford.edu/projects/helix/pubs/process-model/. Mor Peleg, Irene S. Gabashvili, Russ B. Altman |
Proc. IEEE | 3 |
| 2001 | Challenges for knowledge discovery in biologyabstractBioinformatics is the study of information flow in biology. Interest in the field has exploded in the last 10 years with the emergence of techniques for large scale experimental data collection-including genome sequencing, gene expression analysis, protein interaction detection, high-throughput structure determination and others. These techniques, in the context of a large online published literature, have created relatively large data sets (at least by biological standards) that are not possible to analyze manually. There is therefore a critical need for methods to analyze these data and reduce them to new knowledge. The principle challenges to the field include the great diversity of data types and questions that are asked of the data, and the communication difficulties that can exist between experts in biology and experts in machine learning. In this talk, I will provide an introduction to the major biological questions that are being addressed, why they are important, and how the field is trying to address them with technical approaches. Russ B. Altman |
KDD | 1 |
| 2001 | Missing value estimation methods for DNA microarraysabstractMOTIVATION: Gene expression microarray experiments can generate data sets with multiple missing expression values. Unfortunately, many algorithms for gene expression analysis require a complete matrix of gene array values as input. For example, methods such as hierarchical clustering and K-means clustering are not robust to missing data, and may lose effectiveness even with a few missing values. Methods for imputing missing data are needed, therefore, to minimize the effect of incomplete data sets on analyses, and to increase the range of data sets to which these algorithms can be applied. In this report, we investigate automated methods for estimating missing data. RESULTS: We present a comparative study of several methods for the estimation of missing values in gene microarray data. We implemented and evaluated three methods: a Singular Value Decomposition (SVD) based method (SVDimpute), weighted K-nearest neighbors (KNNimpute), and row average. We evaluated the methods using a variety of parameter settings and over different real data sets, and assessed the robustness of the imputation methods to the amount of missing data over the range of 1--20% missing values. We show that KNNimpute appears to provide a more robust and sensitive method for missing value estimation than SVDimpute, and both SVDimpute and KNNimpute surpass the commonly used row average method (as well as filling missing values with zeros). We report results of the comparative experiments and provide recommendations and tools for accurate estimation of missing microarray data under a variety of conditions. Olga G. Troyanskaya, Michael N. Cantor, Gavin Sherlock, Patrick O. Brown, Trevor J. Hastie, Robert Tibshirani, David Botstein, Russ B. Altman |
Bioinform. | 8 |
| 2000 | The new peer review
Isaac S. Kohane, Russ B. Altman |
AMIA | 2 |
| 2000 | Pattern Recognition of Genomic Features with Microarrays: Site Typing of Mycobacterium Tuberculosis Strains
Soumya Raychaudhuri, Joshua M. Stuart, Xuemin Liu, Peter M. Small, Russ B. Altman |
ISMB | 5 |
| 2000 | Viewpoint: The Interactions Between Clinical Informatics and Bioinformatics: A Case StudyabstractFor the past decade, Stanford Medical Informatics has combined clinical informatics and bioinformatics research and training in an explicit way. The interest in applying informatics techniques to both clinical problems and problems in basic science can be traced to the Dendral project in the 1960s. Having bioinformatics and clinical informatics in the same academic unit is still somewhat unusual and can lead to clashes of clinical and basic science cultures. Nevertheless, the benefits of this organization have recently become clear, as the landscape of academic medicine in the next decades has begun to emerge. The author provides examples of technology transfer between clinical informatics and bioinformatics that illustrate how they complement each other. Russ B. Altman |
J. Am. Medical Informatics Assoc. | 1 |
| 1999 | Using imperfect secondary structure predictions to improve molecular structure computationsabstractMOTIVATION: Until ab initio structure prediction methods are perfected, the estimation of structure for protein molecules will depend on combining multiple sources of experimental and theoretical data. Secondary structure predictions are a particularly useful source of structural information, but are currently only approximately 70% correct, on average. Structure computation algorithms which incorporate secondary structure information must therefore have methods for dealing with predictions that are imperfect. EXPERIMENTS PERFORMED: We have modified our algorithm for probabilistic least squares structural computations to accept 'disjunctive' constraints, in which a constraint is provided as a set of possible values, each weighted with a probability. Thus, when a helix is predicted, the distances associated with a helix are given most of the weight, but some weights can be allocated to the other possibilities (strand and coil). We have tested a variety of strategies for this weighting scheme in conjunction with a baseline synthetic set of sparse distance data, and compared it with strategies which do not use disjunctive constraints. RESULTS: Naive interpretations in which predictions were taken as 100% correct led to poor-quality structures. Interpretations that allow disjunctive constraints are quite robust, and even relatively poor predictions (58% correct) can significantly increase the quality of computed structures (almost halving the RMS error from the known structure). CONCLUSIONS: Secondary structure predictions can be used to improve the quality of three-dimensional structural computations. In fact, when interpreted appropriately, imperfect predictions can provide almost as much improvement as perfect predictions in three-dimensional structure calculations. Cheng Che Chen, Jaswinder Pal Singh, Russ B. Altman |
Bioinform. | 3 |
| 1999 | Model Formulation: Automated Diagnosis of Data-Model Conflicts Using MetadataabstractThe authors describe a methodology for helping computational biologists diagnose discrepancies they encounter between experimental data and the predictions of scientific models. The authors call these discrepancies data-model conflicts. They have built a prototype system to help scientists resolve these conflicts in a more systematic, evidence-based manner. In computational biology, data-model conflicts are the result of complex computations in which data and models are transformed and evaluated. Increasingly, the data, models, and tools employed in these computations come from diverse and distributed resources, contributing to a widening gap between the scientist and the original context in which these resources were produced. This contextual rift can contribute to the misuse of scientific data or tools and amplifies the problem of diagnosing data-model conflicts. The authors' hypothesis is that systematic collection of metadata about a computational process can help bridge the contextual rift and provide information for supporting automated diagnosis of these conflicts. The methodology involves three major steps. First, the authors decompose the data-model evaluation process into abstract functional components. Next, they use this process decomposition to enumerate the possible causes of the data-model conflict and direct the acquisition of diagnostically relevant metadata. Finally, they use evidence statically and dynamically generated from the metadata collected to identify the most likely causes of the given conflict. They describe how these methods are implemented in a knowledge-based system called GRENDEL and show how GRENDEL can be used to help diagnose conflicts between experimental data and computationally built structural models of the 30S ribosomal subunit. Richard O. Chen, Russ B. Altman |
J. Am. Medical Informatics Assoc. | 2 |
| 1998 | Bioinformatics in support of molecular medicine
Russ B. Altman |
AMIA | 1 |
| 1998 | MHCWeb: converting a WWW database into a knowledge-based collaborative environment
Lawrence S. Hon, Neil F. Abernethy, Vladimir Brusic, Jenny Chai, Russ B. Altman |
AMIA | 5 |
| 1998 | Updating a bibliography using the related articles function within PubMed
Xuemin Liu, Russ B. Altman |
AMIA | 2 |
| 1998 | A Surface Measure for Probabilistic Structural Computations
Jeanette P. Schmidt, Cheng Che Chen, Jonathan L. Cooper, Russ B. Altman |
ISMB | 4 |
| 1998 | The hierarchical organization of molecular structure computationsabstractThe task of computing molecular structure from combinations of experimental and theoretical constraints is expensive because of the large number of estimated parameters (the 3D coordinates of each atom) and the rugged landscape of many objective functions. For large molecular ensembles with multiple protein and nucleic acid components, the problem of maintaining tractability in structural computations becomes critical. A well known strategy for solving difficult problems is divide and conquer. For molecular computations, there are two ways in which problems can be divided: (1) using the natural hierarchy within biological macromolecules (taking advantage of primary sequence, secondary structural subunits and tertiary structural motifs, when they are known), and (2) using the hierarchy that results from analyzing the distribution of structural constraints (providing information about which substructures are constrained to one another). In this paper, we show that these two hierarchies can be complementary and can provide information for efficient decomposition of structural computations. We demonstrate four methods for building such hierarchies---two automated heuristics that use both natural and empirical hierarchies, one knowledge-based process using both hierarchies, and one method based on the natural hierarchy alone---and apply them to a data set for the procaryotic 30S ribosomal subunit using our probabilistic least squares structure estimation algorithm. We show that the three methods that combine natural hierarchies with empirical hierarchies create decompositions which increase the empirical efficiency of computations by as much as 50-fold. There is only half this gain when using the natural decomposition alone. Although the knowledge-based method performs margi... Cheng Che Chen, Jaswinder Pal Singh, Russ B. Altman |
RECOMB | 3 |
| 1998 | A curriculum for bioinformatics: the time is ripe
Russ B. Altman |
Bioinform. | 1 |
| 1998 | Reuse, CORBA, and knowledge-based systemsabstractBy applying recent advances in the standards for distributed computing, we have developed an architecture for a CORBA implementation of a library of platform-independent, sharable problem-solving methods and knowledge bases. The aim of this library is to allow developers to reuse these components across different tasks and domains. Reuse should be cost-effective; therefore, the library will include standard problem-solving methods whose semantics are well understood and are described with a language for stating the requirements and capabilities of a component. In addition, when a developer needs to adapt a component to a new task, the adaptation costs should be minimal. Thus, we advocate the use of separate mediating components that isolate these adaptations from the original component. We demonstrate our approach with an example: an implementation of a problem-solving method, a knowledge-base server, and mediating components that adapt the method to different knowledge bases and tasks. John H. Gennari, Heyning Cheng, Russ B. Altman, Mark A. Musen |
Int. J. Hum. Comput. Stud. | 3 |
| 1997 | Standardized Representations of the Literature: Combining Diverse Sources of Ribosomal Data
Russ B. Altman, Neil F. Abernethy, Richard O. Chen |
ISMB | 1 |
| 1997 | RIBOWEB: Linking Structural Computations to a Knowledge Base of Published Experimental Data
Richard O. Chen, Ramon M. Felciano, Russ B. Altman |
ISMB | 3 |
| 1996 | Parallel Hierarchical Molecular Structure EstimationabstractDetermining the structure of biological macromolecules such as proteins and nucleic acids is an important element of molecular biology because of the intimate relation between form and function of these molecules. Individual sources of data about molecular structure are subject to varying degrees of uncertainty. Previously we have examined the parallelization of a probabilistic algorithm for combining multiple sources of uncertain data to estimate the three-dimensional structure of molecules and also predict a measure of the uncertainty in the estimated structure. In this paper we extend our work on two major fronts. First we present a hierarchiacal decomposition of the original algorithm which reduces the sequential computational complexity tremendously. The hierarchical decomposition in turn reveals a new axis of parallelism not present in the "flat" organization of the problems, as well as new parallelization problems. We demonstrate good speedups on two cache-coherent shared-memory multiprocessors, the Stanford DASH and the SGI Challenge, with distributed and centralized memory organization, respectively. Our results point to several areas of further study to make both the hierarchiacal and the parallel aspects more flexible for general problems: automatic structure decomposition, processor load balancing across the hierarchy, and data locality management in conjunction with load balancing. Finally we outline the directions we are investigating to incorporate these extensions. Cheng Che Chen, Jaswinder Pal Singh, Russ B. Altman |
SC | 3 |
| 1996 | Constraining volume by matching the moments of a distance distributionabstractThe problem of computing a molecular structure from a set of distances arises in the interpretation of NMR data as well as other experimental methods that yield distance information. Techniques for computing structures must find conformations consistent with the distance data. There are often other constraints on the structure that must be satisfied as well. One of the most problematic constraints is the constraint on the total volume occupied by the atoms. In this paper, we use the first two moments (mean and variance) of an estimated distance distribution to constrain the volume of a computed structure. We show that a probabilistic algorithm for matching the first two moments of the estimated distance distribution significantly improves the quality of the solution, especially when the distance information alone is not sufficient to define the structure precisely. We also show that our method is not sensitive to small errors in the estimates of mean and variance of the distance distribution. Finally, we demonstrate the use of this constraint in computing a low-resolution structure of the 30S prokaryotic ribosomal subunit. Quantitative analysis of our results allows us to assess the information content contained in constraints on volume, and to show that in some cases addition of a volume constraint adds information roughly equivalent to doubling the number of input distances. Our results also demonstrate the flexibility of probabilistic representations of structural constraints, and the importance of including volume information to constrain structural computations-especially in the case of sparse data. Cheng Che Chen, Richard O. Chen, Russ B. Altman |
Comput. Appl. Biosci. | 3 |
| 1995 | Characterizing Oriented Protein Structural Sites Using Biochemical Properties
Steven C. Bagley, Liping Wei, Carol Cheng, Russ B. Altman |
ISMB | 4 |
| 1995 | Using a measure of structural variation to define a core for the globinsabstractAs the database of three-dimensional protein structures expands, it becomes possible to classify related structures into families. Some of these families, such as the globins, have enough members to allow statistical analysis of conserved features. Previously, we have shown that a probabilistic representation based on means and variances can be useful for defining structural cores for large families. These cores contain the subset of atoms that are in essentially the same relative positions in all members of the family. In addition to defining a core, our method creates an ordered list of atoms, ranked by their structural variation. In applying our core-finding procedure to the globins, we find that helices A, B, G and H form a structural core with low variance. These helices fold early in the folding pathway, and superimpose well with helices in the helix-turn-helix repressor protein family. The non-core helices (F and the parts of other helices that interact with it) are associated with the functional differences among the globins, and are encoded within a separate exon. We have also compared the variability measure implicit in our core structures with measures of sequence variability, using a procedure for measuring sequence variability that helps correct for the biased sampling in the databanks. We find, somewhat surprisingly, that sequence variation does not appear to correlate with structural variation. Mark Gerstein, Russ B. Altman |
Comput. Appl. Biosci. | 2 |
| 1995 | A probabilistic approach to determining biological structure: integrating uncertain data sourcesabstractModeling the structure of biological molecules is critical for understanding how these structures perform their function, and for designing compounds to modify or enhance this function (for medicinal or industrial purposes). The determination of molecular structure involves defining three-dimensional positions for each of the constituent atoms using a variety of experimental, theoretical and empirical data sources. Unfortunately, each of these data sources can be noisy or not available in sufficient abundance to determine the precise position of each atom. Instead, some atomic positions are precisely defined by the data, and others are poorly defined. An understanding of structural uncertainty is critical for properly interpreting structural models. We have developed a Bayesian approach for determining the coordinates of atoms in a three-dimensional space. Our algorithm takes as input a set of probabilistic constraints on the coordinates of the atoms, and an a priori distribution for each atom location. The output is a maximum a posteriori (MAP) estimate of the location of each atom. We introduce constraints as updates to the prior distributions. In this paper, we describe the algorithm and show its performance on three data sets. The first data set is synthetic and illustrates the convergence properties of the method. The other data sets comprise real biological data for a protein (the trp repressor molecule) and a nucleic acid (the transfer RNA fold). Finally, we describe how we have begun to extend the algorithm to make it suitable for non-Gaussian constraints. Russ B. Altman |
Int. J. Hum. Comput. Stud. | 1 |
| 1994 | Finding an Average Core Structure: Application to the Globins
Russ B. Altman, Mark Gerstein |
ISMB | 1 |
| 1994 | Constraint Satisfaction Techniques for Modeling Large Complexes: Application to the Central Domain of 16S Ribosomal RNA
Russ B. Altman, Bryn Weiser, Harry F. Noller |
ISMB | 1 |
| 1994 | Parallel protein structure determination from uncertain dataabstractMolecular structure determination is an important task in biology because of the intimate relation between form and function of biological molecules. Individual sources of information about molecular structure are subject to uncertainty and are not sufficiently abundant to define the structure to high accuracy by themselves. The authors have examined a probabilistic algorithm, PROTEAN, which can incorporate multiple sources of uncertain data to estimate the three dimensional structure of molecules and also predict a measure of the uncertainty in the estimated structure. They have applied this algorithm successfully to several biological structure problems. Like most structure prediction methods, this algorithm is computationally expensive for realistic biological macromolecules. The authors experiment with speeding up the algorithm through the application of parallelism. They present a parallel version of the algorithm, and demonstrate good speedups on a 32-processor Stanford DASH, a cache coherent shared address space multiprocessor. The results were obtained by exploiting data locality only in the per-processor coherent caches, without attempt to distribute data intelligently in the physically distributed main memory of the machine. The authors also obtained very good speedups on a state of the art commercial multiprocessor, the Silicon Graphics Challenge. Finally, the authors propose an extension to the serial algorithm which enables it to handle a wider class of data, and discuss the potential for parallelization of the extended algorithm.> Cheng Che Chen, Jaswinder Pal Singh, William B. Poland, Russ B. Altman |
SC | 4 |
| 1994 | Probabilistic Constraint Satisfaction with Non-Gaussian Noise
Russ B. Altman, Cheng Che Chen, William B. Poland, Jaswinder Pal Singh |
UAI | 1 |
| 1993 | Probabilistic Structure Calculations: A Three-Dimensional tRNA Structure from Sequence Correlation Data
Russ B. Altman |
ISMB | 1 |
| 1993 | A Probabilistic Algorithm for Calculating Structure: Borrowing from Simulated Annealing
Russ B. Altman |
UAI | 1 |
| 1987 | Partial Compilation of Strategic Knowledge
Russ B. Altman, Bruce G. Buchanan |
AAAI | 1 |
| 1986 | PROTEAN: Deriving Protein Structure from Constraints
Barbara Hayes-Roth, Bruce G. Buchanan, Olivier Lichtarge, Mike Hewitt, Russ B. Altman, James F. Brinkley, Craig Cornelius, Bruce S. Duncan, Oleg Jardetzky |
AAAI | 5 |