EDBT 2026 Demo / reviewers in the wild / expert
Michael Krauthammer
dblp:90/4978
· DBLP profile ↗
42ranked-venue papers
7as first author
5since 2021 · last 2025
0000-0002-4808-1845ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 36 · 7 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
10 papers |
Bioinformatics and computational biology · 96% Medical and health informatics · 4% | |
| Artificial intelligence
2 papers |
Generative modeling · 87% Knowledge representation and reasoning · 13% |
Topics — the 28 heaviest of 30, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
flow matching |
0.9 | 1 | 2025 | A Variational Perspective on Generative Protein Fitness Optimization · ICML 2025 |
Bioinformatics and computational biology
protein design |
0.9 | 1 | 2025 | A Variational Perspective on Generative Protein Fitness Optimization · ICML 2025 |
Bioinformatics and computational biology › protein design
protein fitness optimization |
0.9 | 1 | 2025 | A Variational Perspective on Generative Protein Fitness Optimization · ICML 2025 |
Bioinformatics and computational biology › genomics › genomic data analysis
cell-free DNA analysis |
0.8 | 1 | 2024 | Fragmentstein - facilitating data reuse for cell-free DNA fragment analysis · Bioinform. 2024 |
Bioinformatics and computational biology › genomics
variant annotation |
0.3 | 2 | 2015 | Mutadelic: mutation analysis using description logic inferencing capabilities · Bioinform. 2015 MU2A - reconciling the genome and transcriptome to determine the effects of base substitutions · Bioinform. 2011 |
Bioinformatics and computational biology
sequence analysis |
0.3 | 1 | 2017 | PySeqLab: an open source Python package for sequence labeling and segmentation · Bioinform. 2017 |
Bioinformatics and computational biology › sequence analysis › sequence modeling
sequence labeling |
0.3 | 1 | 2017 | PySeqLab: an open source Python package for sequence labeling and segmentation · Bioinform. 2017 |
Bioinformatics and computational biology › sequence analysis › sequence annotation
sequence segmentation |
0.3 | 1 | 2017 | PySeqLab: an open source Python package for sequence labeling and segmentation · Bioinform. 2017 |
Bioinformatics and computational biology › cancer genomics › copy number analysis
allele-specific copy number |
0.2 | 1 | 2016 | Global copy number profiling of cancer genomes · Bioinform. 2016 |
Bioinformatics and computational biology
cancer genomics |
0.2 | 1 | 2016 | Global copy number profiling of cancer genomes · Bioinform. 2016 |
Bioinformatics and computational biology
genomic privacy |
0.2 | 1 | 2024 | Fragmentstein - facilitating data reuse for cell-free DNA fragment analysis · Bioinform. 2024 |
Medical and health informatics
clinical genomics |
0.2 | 1 | 2015 | Mutadelic: mutation analysis using description logic inferencing capabilities · Bioinform. 2015 |
Bioinformatics and computational biology › clinical bioinformatics
variant interpretation |
0.2 | 1 | 2015 | Mutadelic: mutation analysis using description logic inferencing capabilities · Bioinform. 2015 |
Information retrieval
search engines |
0.1 | 1 | 2008 | Yale Image Finder (YIF): a new search engine for retrieving biomedical images · Bioinform. 2008 |
Multimedia analysis and retrieval › image retrieval
content-based image retrieval |
0.1 | 1 | 2008 | Yale Image Finder (YIF): a new search engine for retrieving biomedical images · Bioinform. 2008 |
Multimedia analysis and retrieval
image retrieval |
0.1 | 1 | 2008 | Yale Image Finder (YIF): a new search engine for retrieving biomedical images · Bioinform. 2008 |
Bioinformatics and computational biology › cancer genomics
tumor purity estimation |
0.1 | 1 | 2016 | Global copy number profiling of cancer genomes · Bioinform. 2016 |
Bioinformatics and computational biology
proteomics |
0.1 | 1 | 2007 | Leveraging the structure of the Semantic Web to enhance information retrieval for proteomics · Bioinform. 2007 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
abductive reasoning |
0.1 | 1 | 2015 | Mutadelic: mutation analysis using description logic inferencing capabilities · Bioinform. 2015 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
description logic |
0.1 | 1 | 2015 | Mutadelic: mutation analysis using description logic inferencing capabilities · Bioinform. 2015 |
Bioinformatics and computational biology › network bioinformatics › biological network analysis
biological network inference |
0.0 | 1 | 2004 | Probabilistic inference of molecular networks from noisy data sources · Bioinform. 2004 |
Bioinformatics and computational biology
protein-protein interaction prediction |
0.0 | 1 | 2004 | Probabilistic inference of molecular networks from noisy data sources · Bioinform. 2004 |
Bioinformatics and computational biology › biological network › network biology
molecular interaction network |
0.0 | 1 | 2002 | Of truth and pathways: chasing bits of information through myriads of articles · ISMB 2002 |
Bioinformatics and computational biology › systems biology
gene regulatory network modeling |
0.0 | 1 | 2000 | A knowledge model for analysis and simulation of regulatory networks · Bioinform. 2000 |
Bioinformatics and computational biology › ontology
ontology development |
0.0 | 1 | 2000 | A knowledge model for analysis and simulation of regulatory networks · Bioinform. 2000 |
Information retrieval › query reformulation
query expansion |
0.0 | 1 | 2007 | Leveraging the structure of the Semantic Web to enhance information retrieval for proteomics · Bioinform. 2007 |
Graph data management
RDF graph |
0.0 | 1 | 2007 | Leveraging the structure of the Semantic Web to enhance information retrieval for proteomics · Bioinform. 2007 |
Bioinformatics and computational biology
signal transduction |
0.0 | 1 | 2000 | A knowledge model for analysis and simulation of regulatory networks · Bioinform. 2000 |
Methods — techniques the papers use, named apart from their topics
variational inference · 1.7flow matching · 1.7fitness predictor · 1.7reference genome sequence completion · 0.8viterbi algorithm · 0.3structured perceptron · 0.3stochastic gradient descent · 0.3conditional random field · 0.3log r ratio analysis · 0.2b-allele frequency analysis · 0.2description logic inferencing · 0.2abductive reasoning · 0.2optical character recognition · 0.2image content similarity · 0.2inverse document frequency · 0.1cosine similarity · 0.1RDF graph · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Variational Perspective on Generative Protein Fitness OptimizationabstractThe goal of protein fitness optimization is to discover new protein variants with enhanced fitness for a given use. The vast search space and the sparsely populated fitness landscape, along with the discrete nature of protein sequences, pose significant challenges when trying to determine the gradient towards configurations with higher fitness. We introduce *Variational Latent Generative Protein Optimization* (VLGPO), a variational perspective on fitness optimization. Our method embeds protein sequences in a continuous latent space to enable efficient sampling from the fitness distribution and combines a (learned) flow matching prior over sequence mutations with a fitness predictor to guide optimization towards sequences with high fitness. VLGPO achieves state-of-the-art results on two different protein benchmarks of varying complexity. Moreover, the variational design with explicit prior and likelihood functions offers a flexible plug-and-play framework that can be easily customized to suit various protein design tasks. Lea Bogensperger, Dominik Narnhofer, Konrad Schindler, Michael Krauthammer |
ICML | 5 |
| 2025 | Deep learning for automated organ activity profiling in sarcoidosisabstractAbstract Background Sarcoidosis is a complex, multisystem inflammatory disease with variable organ involvement and course1. 18F-FDG PET/CT is the gold standard for activity assessment but is costly, scarce, and operator-dependent. Methods We present a fully automated deep-learning pipeline for multi-organ PET/CT quantification, applied to 180 patients with 770 longitudinal scans linked to clinical and laboratory data. Based on contrast-enhanced CT, the lungs, myocardium, spleen, and liver are segmented using TotalSegmentator2; co-registered PET images provide organ-level SUVmax values and metabolic activity scores. Results Pipeline-derived metrics showed strong agreement with clinician assessments of cardiac involvement (for activity, r = 0.93; for SUVmax, r = 0.95, both p < 0.001). Organ-resolved analyses revealed significant correlations between PET metrics and inflammatory biomarkers, particularly in lung involvement (e.g., serum TNF-α, r = 0.39, FDR < 0.05). Using repeated measurements of disease activity and biomarker concentrations, longitudinal modeling captured flare–remission dynamics. Biomarker trajectories paralleled PET activity, supporting their potential as non-invasive surrogate markers when PET/CT is unavailable; in practice, routine lab tests between imaging visits could flag incoming flares and guide early treatment adjustments. Conclusion This automated, organ-resolved PET/CT pipeline, paired with routine labs, enables objective, scalable sarcoidosis profiling beyond tertiary clinics, facilitating early and precise treatment response monitoring and prognostication. References 1. Grunewald J., Grutters J.C., Arkema E.V., Saketkoo L.A., Moller D.R., Müller-Quernheim J. ‘Sarcoidosis.’ Nature Reviews Disease Primers 2019;5(1):45. 2. Wasserthal J., Breit H.C., Meyer M.T., Pradella M., Hinck D., et al. ‘TotalSegmentator: robust segmentation of 104 anatomic structures in CT images.’ Radiology: Artificial Intelligence 2023;5(5):e230024. Sonja Katz, Bowen Fan, Michael Krauthammer, Jakob Nilsson |
Briefings Bioinform. | 3 |
| 2024 | Simple Contrastive Representation Learning for Time Series ForecastingabstractContrastive learning methods have shown an impressive ability to learn meaningful representations for image or time series classification. However, these methods are less effective for time series forecasting, as optimization of instance discrimination is not directly applicable to predicting the future state from the historical context. To address these limitations, we propose SimTS, a simple representation learning approach for improving time series forecasting by learning to predict the future from the past in the latent space. SimTS exclusively uses positive pairs and does not depend on negative pairs or specific characteristics of a given time series. In addition, we show the shortcomings of the current contrastive learning framework used for time series forecasting through a detailed ablation study. Overall, our work suggests that SimTS is a promising alternative to other contrastive learning approaches for time series forecasting. Manuel Schürch, Amina Mollaysa, Michael Krauthammer |
ICASSP | 6 |
| 2024 | Fragmentstein - facilitating data reuse for cell-free DNA fragment analysisabstractSUMMARY: Method development for the analysis of cell-free DNA (cfDNA) sequencing data is impeded by limited data sharing due to the strict control of sensitive genomic data. An existing solution for facilitating data sharing removes nucleotide-level information from raw cfDNA sequencing data, keeping alignment coordinates only. This simplified format can be publicly shared and would, theoretically, suffice for common functional analyses of cfDNA data. However, current bioinformatics software requires nucleotide-level information and cannot process the simplified format. We present Fragmentstein, a command-line tool for converting non-sensitive cfDNA-fragmentation data into alignment mapping (BAM) files. Fragmentstein complements fragment coordinates with sequence information from a reference genome to reconstruct BAM files. We demonstrate the utility of Fragmentstein by showing the feasibility of copy number variant (CNV), nucleosome occupancy, and fragment length analyses from non-sensitive fragmentation data. AVAILABILITY AND IMPLEMENTATION: Implemented in bash, Fragmentstein is available at https://github.com/uzh-dqbm-cmi/fragmentstein, licensed under GNU GPLv3. Zsolt Balázs, Todor Gitchev, Ivna Ivankovic, Michael Krauthammer |
Bioinform. | 4 |
| 2021 | AttentionDDI: Siamese attention-based deep learning method for drug-drug interaction predictionsabstractBACKGROUND: Drug-drug interactions (DDIs) refer to processes triggered by the administration of two or more drugs leading to side effects beyond those observed when drugs are administered by themselves. Due to the massive number of possible drug pairs, it is nearly impossible to experimentally test all combinations and discover previously unobserved side effects. Therefore, machine learning based methods are being used to address this issue. METHODS: We propose a Siamese self-attention multi-modal neural network for DDI prediction that integrates multiple drug similarity measures that have been derived from a comparison of drug characteristics including drug targets, pathways and gene expression profiles. RESULTS: Our proposed DDI prediction model provides multiple advantages: (1) It is trained end-to-end, overcoming limitations of models composed of multiple separate steps, (2) it offers model explainability via an Attention mechanism for identifying salient input features and (3) it achieves similar or better prediction performance (AUPR scores ranging from 0.77 to 0.92) compared to state-of-the-art DDI models when tested on various benchmark datasets. Novel DDI predictions are further validated using independent data resources. CONCLUSIONS: We find that a Siamese multi-modal neural network is able to accurately predict DDIs and that an Attention mechanism, typically used in the Natural Language Processing domain, can be beneficially applied to aid in DDI model explainability. Kyriakos Schwarz, Nicolas Andres Perez Gonzalez, Michael Krauthammer |
BMC Bioinform. | 4 |
| 2020 | Iterations for Propensity Score Matching in MonetDB
Michael H. Böhlen, Oksana Dolmatova, Michael Krauthammer, Alphonse Mariyagnanaseelan, Jonathan Stahl, Timo Surbeck |
ADBIS | 3 |
| 2018 | Prototyping a precision oncology 3.0 rapid learning platformabstractBACKGROUND: We describe a prototype implementation of a platform that could underlie a Precision Oncology Rapid Learning system. RESULTS: We describe the prototype platform, and examine some important issues and details. In the Appendix we provide a complete walk-through of the prototype platform. CONCLUSIONS: The design choices made in this implementation rest upon ten constitutive hypotheses, which, taken together, define a particular view of how a rapid learning medical platform might be defined, organized, and implemented. Connor Sweetnam, Simone Mocellin, Michael Krauthammer, Nathaniel Knopf, Robert Baertsch, Jeff Shrager |
BMC Bioinform. | 3 |
| 2017 | PySeqLab: an open source Python package for sequence labeling and segmentationabstractMOTIVATION: Text and genomic data are composed of sequential tokens, such as words and nucleotides that give rise to higher order syntactic constructs. In this work, we aim at providing a comprehensive Python library implementing conditional random fields (CRFs), a class of probabilistic graphical models, for robust prediction of these constructs from sequential data. RESULTS: Python Sequence Labeling (PySeqLab) is an open source package for performing supervised learning in structured prediction tasks. It implements CRFs models, that is discriminative models from (i) first-order to higher-order linear-chain CRFs, and from (ii) first-order to higher-order semi-Markov CRFs (semi-CRFs). Moreover, it provides multiple learning algorithms for estimating model parameters such as (i) stochastic gradient descent (SGD) and its multiple variations, (ii) structured perceptron with multiple averaging schemes supporting exact and inexact search using 'violation-fixing' framework, (iii) search-based probabilistic online learning algorithm (SAPO) and (iv) an interface for Broyden-Fletcher-Goldfarb-Shanno (BFGS) and the limited-memory BFGS algorithms. Viterbi and Viterbi A* are used for inference and decoding of sequences. Using PySeqLab, we built models (classifiers) and evaluated their performance in three different domains: (i) biomedical Natural language processing (NLP), (ii) predictive DNA sequence analysis and (iii) Human activity recognition (HAR). State-of-the-art performance comparable to machine-learning based systems was achieved in the three domains without feature engineering or the use of knowledge sources. AVAILABILITY AND IMPLEMENTATION: PySeqLab is available through https://bitbucket.org/A_2/pyseqlab with tutorials and documentation. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Michael Krauthammer |
Bioinform. | 2 |
| 2017 | Toward automated assessment of health Web page quality using the DISCERN instrumentabstractBACKGROUND: As the Internet becomes the number one destination for obtaining health-related information, there is an increasing need to identify health Web pages that convey an accurate and current view of medical knowledge. In response, the research community has created multicriteria instruments for reliably assessing online medical information quality. One such instrument is DISCERN, which measures health Web page quality by assessing an array of features. In order to scale up use of the instrument, there is interest in automating the quality evaluation process by building machine learning (ML)-based DISCERN Web page classifiers. OBJECTIVE: The paper addresses 2 key issues that are essential before constructing automated DISCERN classifiers: (1) generation of a robust DISCERN training corpus useful for training classification algorithms, and (2) assessment of the usefulness of the current DISCERN scoring schema as a metric for evaluating the performance of these algorithms. METHODS: Using DISCERN, 272 Web pages discussing treatment options in breast cancer, arthritis, and depression were evaluated and rated by trained coders. First, different consensus models were compared to obtain a robust aggregated rating among the coders, suitable for a DISCERN ML training corpus. Second, a new DISCERN scoring criterion was proposed (features-based score) as an ML performance metric that is more reflective of the score distribution across different DISCERN quality criteria. RESULTS: First, we found that a probabilistic consensus model applied to the DISCERN instrument was robust against noise (random ratings) and superior to other approaches for building a training corpus. Second, we found that the established DISCERN scoring schema (overall score) is ill-suited to measure ML performance for automated classifiers. CONCLUSION: Use of a probabilistic consensus model is advantageous for building a training corpus for the DISCERN instrument, and use of a features-based score is an appropriate ML metric for automated DISCERN classifiers. AVAILABILITY: The code for the probabilistic consensus model is available at https://bitbucket.org/A_2/em_dawid/ . Peter J. Schulz, Michael Krauthammer |
J. Am. Medical Informatics Assoc. | 3 |
| 2016 | Controlling testing volume for respiratory viruses using machine learning and text mining
Mark V. Mai, Michael Krauthammer |
AMIA | 2 |
| 2016 | Global copy number profiling of cancer genomesabstractUNLABELLED: In this article, we introduce a robust and efficient strategy for deriving global and allele-specific copy number alternations (CNA) from cancer whole exome sequencing data based on Log R ratios and B-allele frequencies. Applying the approach to the analysis of over 200 skin cancer samples, we demonstrate its utility for discovering distinct CNA events and for deriving ancillary information such as tumor purity. AVAILABILITY AND IMPLEMENTATION: https://github.com/xfwang/CLOSE CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mengjie Chen, Xiaoqing Yu, Natapol Pornputtapong, Hao Chen 0064, Nancy Ruonan Zhang, R. Scott Powers, Michael Krauthammer |
Bioinform. | 8 |
| 2015 | Publishing Without Publishers: A Decentralized Approach to Dissemination, Retrieval, and Archiving of DataabstractMaking available and archiving scientific results is for the most part still considered the task of classical publishing companies, despite the fact that classical forms of publishing centered around printed narrative articles no longer seem well-suited in the digital age. In particular, there exist currently no efficient, reliable, and agreed-upon methods for publishing scientific datasets, which have become increasingly important for science. Here we propose to design scientific data publishing as a Web-based bottom-up process, without top-down control of central authorities such as publishing companies. Based on a novel combination of existing concepts and technologies, we present a server network to decentrally store and archive data in the form of nanopublications, an RDF-based format to represent scientific data. We show how this approach allows researchers to publish, retrieve, verify, and recombine datasets of nanopublications in a reliable and trustworthy manner, and we argue that this architecture could be used for the Semantic Web in general. Evaluation of the current small network shows that this system is efficient and reliable. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves. Tobias Kuhn, Christine Chichester, Michael Krauthammer, Michel Dumontier |
ISWC (1) | 3 |
| 2015 | Mutadelic: mutation analysis using description logic inferencing capabilitiesabstractMOTIVATION: As next generation sequencing gains a foothold in clinical genetics, there is a need for annotation tools to characterize increasing amounts of patient variant data for identifying clinically relevant mutations. While existing informatics tools provide efficient bulk variant annotations, they often generate excess information that may limit their scalability. RESULTS: We propose an alternative solution based on description logic inferencing to generate workflows that produce only those annotations that will contribute to the interpretation of each variant. Workflows are dynamically generated using a novel abductive reasoning framework called a basic framework for abductive workflow generation (AbFab). Criteria for identifying disease-causing variants in Mendelian blood disorders were identified and implemented as AbFab services. A web application was built allowing users to run workflows generated from the criteria to analyze genomic variants. Significant variants are flagged and explanations provided for why they match or fail to match the criteria. AVAILABILITY AND IMPLEMENTATION: The Mutadelic web application is available for use at http://krauthammerlab.med.yale.edu/mutadelic. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Matthew Holford, Michael Krauthammer |
Bioinform. | 2 |
| 2013 | Nanorecords: A Novel Approach to Communicating High-Level Medical Information
Mark V. Mai, Tobias Kuhn, José Costa, Michael Krauthammer |
AMIA | 4 |
| 2013 | Broadening the Scope of Nanopublications
Tobias Kuhn, Paolo Emilio Barbano, Mate Levente Nagy, Michael Krauthammer |
ESWC | 4 |
| 2013 | Complementary ensemble clustering of biomedical data
Samah Jamal Fodeh, Cynthia Brandt, Thaibinh Luong, Ali Haddad, Martin H. Schultz, Terrence Murphy, Michael Krauthammer |
J. Biomed. Informatics | 7 |
| 2012 | Finding and Accessing Diagrams in Biomedical Publications
Tobias Kuhn, Thaibinh Luong, Michael Krauthammer |
AMIA | 3 |
| 2012 | Estimating a gene's mutation burden by the number of observed synonymous base substitutionsabstractA common goal of tumor sequencing projects is the identification of genes whose mutations are selected for during tumor development. This is accomplished by finding genes that have more nonsynonymous mutations than expected by an estimated background mutation frequency. While this frequency is unknown, it can be estimated using both the observed synonymous mutation frequency, and the nonsynonymous to synonymous mutation ratio. The synonymous mutation frequency can be determined across all genes, or in a gene-specific manner. This choice introduces an interesting tradeoff. A gene-specific frequency is difficult to estimate given small or missing synonymous mutation counts, but adjusts for an underlying mutation load bias. Using a genome-wide synonymous frequency is more robust, but is less suited for adjusting for the same bias. Studying three evaluation criteria for identifying genes with high nonsynonymous mutation burden (preferential selection of expressed genes, genes with mutations in conserved bases, and genes that show loss of heterozygosity), we find that the gene-specific synonymous frequency is superior in the gene expression and conservation tests, while both frequencies perform similarly for the loss of heterozygosity test. In conclusion, we believe that the use of the gene-specific synonymous mutation frequency is well suited for estimating a gene's nonsynonymous mutation burden. Perry Evans, Michael Krauthammer |
BIBM | 2 |
| 2012 | A semantic web framework to integrate cancer omics data with biological knowledgeabstractBACKGROUND: The RDF triple provides a simple linguistic means of describing limitless types of information. Triples can be flexibly combined into a unified data source we call a semantic model. Semantic models open new possibilities for the integration of variegated biological data. We use Semantic Web technology to explicate high throughput clinical data in the context of fundamental biological knowledge. We have extended Corvus, a data warehouse which provides a uniform interface to various forms of Omics data, by providing a SPARQL endpoint. With the querying and reasoning tools made possible by the Semantic Web, we were able to explore quantitative semantic models retrieved from Corvus in the light of systematic biological knowledge. RESULTS: For this paper, we merged semantic models containing genomic, transcriptomic and epigenomic data from melanoma samples with two semantic models of functional data - one containing Gene Ontology (GO) data, the other, regulatory networks constructed from transcription factor binding information. These two semantic models were created in an ad hoc manner but support a common interface for integration with the quantitative semantic models. Such combined semantic models allow us to pose significant translational medicine questions. Here, we study the interplay between a cell's molecular state and its response to anti-cancer therapy by exploring the resistance of cancer cells to Decitabine, a demethylating agent. CONCLUSIONS: We were able to generate a testable hypothesis to explain how Decitabine fights cancer - namely, that it targets apoptosis-related gene promoters predominantly in Decitabine-sensitive cell lines, thus conveying its cytotoxic effect by activating the apoptosis pathway. Our research provides a framework whereby similar hypotheses can be developed easily. Matthew Holford, Jamie P. McCusker, Kei-Hoi Cheung, Michael Krauthammer |
BMC Bioinform. | 4 |
| 2011 | MU2A - reconciling the genome and transcriptome to determine the effects of base substitutionsabstractMOTIVATION: Next-generation sequencing technologies enable the identification of sequence variation in the genome and transcriptome. Differences between the reference genome and transcript libraries complicate the determination of the effect of genomic sequence variants on protein products; similarly, these differences complicate the mapping of sequence variants found in transcripts to their respective genomic position. We have developed MU2A, a publicly available web service for variant annotation that reconciles differences between the genome and transcriptome, enabling the rapid and accurate determination of the effects of genomic variants on protein products, and the mapping of variants detected in transcripts to genomic coordinates. The MU2A web service is available at http://krauthammerlab.med.yale.edu/mu2a. We have released MU2A as open source, available at http://code.google.com/p/mu2a/. Vijay Garla, Sebastian Szpakowski, Michael Krauthammer |
Bioinform. | 4 |
| 2010 | A new pivoting and iterative text detection algorithm for biomedical images
Songhua Xu, Michael Krauthammer |
J. Biomed. Informatics | 2 |
| 2009 | Enriching PubMed Related Article Search with Sentence Level Co-citations
Nam Tran, Pedro Alves, Shuangge Ma, Michael Krauthammer |
AMIA | 4 |
| 2009 | Embedding the Guideline Elements Model in Web Ontology Language
Nam Tran, George Michel, Michael Krauthammer, Richard N. Shiffman |
AMIA | 3 |
| 2009 | Semantic web data warehousing for caGridabstractThe National Cancer Institute (NCI) is developing caGrid as a means for sharing cancer-related data and services. As more data sets become available on caGrid, we need effective ways of accessing and integrating this information. Although the data models exposed on caGrid are semantically well annotated, it is currently up to the caGrid client to infer relationships between the different models and their classes. In this paper, we present a Semantic Web-based data warehouse (Corvus) for creating relationships among caGrid models. This is accomplished through the transformation of semantically-annotated caBIG Unified Modeling Language (UML) information models into Web Ontology Language (OWL) ontologies that preserve those semantics. We demonstrate the validity of the approach by Semantic Extraction, Transformation and Loading (SETL) of data from two caGrid data sources, caTissue and caArray, as well as alignment and query of those sources in Corvus. We argue that semantic integration is necessary for integration of data from distributed web services and that Corvus is a useful way of accomplishing this. Our approach is generalizable and of broad utility to researchers facing similar integration challenges. Jamie P. McCusker, Joshua A. Phillips, Alejandra N. González-Beltrán, Anthony Finkelstein, Michael Krauthammer |
BMC Bioinform. | 5 |
| 2009 | Structural similarity assessment for drug sensitivity prediction in cancerabstractBACKGROUND: The ability to predict drug sensitivity in cancer is one of the exciting promises of pharmacogenomic research. Several groups have demonstrated the ability to predict drug sensitivity by integrating chemo-sensitivity data and associated gene expression measurements from large anti-cancer drug screens such as NCI-60. The general approach is based on comparing gene expression measurements from sensitive and resistant cancer cell lines and deriving drug sensitivity profiles consisting of lists of genes whose expression is predictive of response to a drug. Importantly, it has been shown that such profiles are generic and can be applied to cancer cell lines that are not part of the anti-cancer screen. However, one limitation is that the profiles can not be generated for untested drugs (i.e., drugs that are not part of an anti-cancer drug screen). In this work, we propose using an existing drug sensitivity profile for drug A as a substitute for an untested drug B given high structural similarities between drugs A and B. RESULTS: We first show that structural similarity between pairs of compounds in the NCI-60 dataset highly correlates with the similarity between their activities across the cancer cell lines. This result shows that structurally similar drugs can be expected to have a similar effect on cancer cell lines. We next set out to test our hypothesis that we can use existing drug sensitivity profiles as substitute profiles for untested drugs. In a cross-validation experiment, we found that the use of substitute profiles is possible without a significant loss of prediction accuracy if the substitute profile was generated from a compound with high structural similarity to the untested compound. CONCLUSION: Anti-cancer drug screens are a valuable resource for generating omics-based drug sensitivity profiles. We show that it is possible to extend the usefulness of existing screens to untested drugs by deriving substitute sensitivity profiles from structurally similar drugs part of the screen. Pavithra Shivakumar, Michael Krauthammer |
BMC Bioinform. | 2 |
| 2008 | Yale Image Finder (YIF): a new search engine for retrieving biomedical imagesabstractUNLABELLED: Yale Image Finder (YIF) is a publicly accessible search engine featuring a new way of retrieving biomedical images and associated papers based on the text carried inside the images. Image queries can also be issued against the image caption, as well as words in the associated paper abstract and title. A typical search scenario using YIF is as follows: a user provides few search keywords and the most relevant images are returned and presented in the form of thumbnails. Users can click on the image of interest to retrieve the high resolution image. In addition, the search engine will provide two types of related images: those that appear in the same paper, and those from other papers with similar image content. Retrieved images link back to their source papers, allowing users to find related papers starting with an image of interest. Currently, YIF has indexed over 140 000 images from over 34 000 open access biomedical journal papers. AVAILABILITY: http://krauthammerlab.med.yale.edu/imagefinder/ Songhua Xu, Jamie P. McCusker, Michael Krauthammer |
Bioinform. | 3 |
| 2007 | Leveraging the structure of the Semantic Web to enhance information retrieval for proteomicsabstractMOTIVATION: Proteomics researchers need to be able to quickly retrieve relevant information from the web and the biomedical literature. To improve information retrieval, we leverage the structure of the semantic web, developing an approach for joining it with the largely opposing paradigm of unsupervised web search. RESULTS: Our approach uses a Resource-Description-Framework (RDF) graph that inter-relates documents through their associated biological identifiers (e.g., protein ID). A search begins with a simple query term (UniProt identifier), which is expanded with terms extracted from documents in the RDF graph surrounding the query ("the subgraph"). We re-rank documents in the full corpus (e.g. all PubMed) by their cosine-similarity scores against a composite word-weight vector created from the subgraph. This vector is a weighted sum of individual word-weight vectors for documents at each node of the subgraph, taking into account the types of relationships between the central query identifier and the nodes connected to it. The computation also uses inverse document frequency (IDF) in a novel way to rescale the local word frequencies in the query's subgraph relative to that in other subgraphs. Applying our procedure to PubMed, we optimize weights for various relationships in the subgraph and benchmark overall performance in detail. Using a subgraph containing family relationships (from PFAM) results in a significant improvement in accuracy (as compared to not considering the subgraph in the search) when assessed against known relationships in the yeast literature. Moreover, we achieve this accuracy using only relatively simple and computationally efficient methods. Andrew K. Smith, Kei-Hoi Cheung, Michael Krauthammer, Martin H. Schultz, Mark Gerstein |
Bioinform. | 3 |
| 2006 | Shallow Semantic Parsing of Randomized Controlled Trial Reports
Hyung M. Paek, Yacov Kogan, Prem Thomas, Seymour Codish, Michael Krauthammer |
AMIA | 5 |
| 2006 | A semantic web approach to biological pathway data reasoning and integration
Kei-Hoi Cheung, Peishen Qi, David Tuck, Michael Krauthammer |
J. Web Semant. | 4 |
| 2005 | Towards Semantic Role Labeling & IE in the Medical Literature
Yacov Kogan, Nigel Collier, Serguei V. S. Pakhomov, Michael Krauthammer |
AMIA | 4 |
| 2004 | Probabilistic inference of molecular networks from noisy data sourcesabstractInformation on molecular networks, such as networks of interacting proteins, comes from diverse sources that contain remarkable differences in distribution and quantity of errors. Here, we introduce a probabilistic model useful for predicting protein interactions from heterogeneous data sources. The model describes stochastic generation of protein-protein interaction networks with real-world properties, as well as generation of two heterogeneous sources of protein-interaction information: research results automatically extracted from the literature and yeast two-hybrid experiments. Based on the domain composition of proteins, we use the model to predict protein interactions for pairs of proteins for which no experimental data are available. We further explore the prediction limits, given experimental data that cover only part of the underlying protein networks. This approach can be extended naturally to include other types of biological data sources. Ivan Iossifov, Michael Krauthammer, Carol Friedman, Vasileios Hatzivassiloglou, Joel S. Bader, Kevin P. White, Andrey Rzhetsky |
Bioinform. | 2 |
| 2004 | Term identification in the biomedical literature
Michael Krauthammer, Goran Nenadic |
J. Biomed. Informatics | 1 |
| 2004 | GeneWays: a system for extracting, analyzing, visualizing, and integrating molecular pathway data
Andrey Rzhetsky, Ivan Iossifov, Tomohiro Koike, Michael Krauthammer, Pauline Kra, Mitzi Morris, Hong Yu 0001, Pablo Ariel Duboue, Wubin Weng, W. John Wilbur |
J. Biomed. Informatics | 4 |
| 2003 | A Native XML Database Design for Clinical Document Research
Stephen B. Johnson, David A. Campbell, Michael Krauthammer, P. Karina Tulipano, Eneida A. Mendonça, Carol Friedman, George Hripcsak |
AMIA | 3 |
| 2002 | Representing nested semantic information in a linear string of text using XML
Michael Krauthammer, Stephen B. Johnson, George Hripcsak, David A. Campbell, Carol Friedman |
AMIA | 1 |
| 2002 | Of truth and pathways: chasing bits of information through myriads of articlesabstractKnowledge on interactions between molecules in living cells is indispensable for theoretical analysis and practical applications in modern genomics and molecular biology. Building such networks relies on the assumption that the correct molecular interactions are known or can be identified by reading a few research articles. However, this assumption does not necessarily hold, as truth is rather an emerging property based on many potentially conflicting facts. This paper explores the processes of knowledge generation and publishing in the molecular biology literature using modelling and analysis of real molecular interaction data. The data analysed in this article were automatically extracted from 50000 research articles in molecular biology using a computer system called GeneWays containing a natural language processing module. The paper indicates that truthfulness of statements is associated in the minds of scientists with the relative importance (connectedness) of substances under study, revealing a potential selection bias in the reporting of research results. Aiming at understanding the statistical properties of the life cycle of biological facts reported in research articles, we formulate a stochastic model describing generation and propagation of knowledge about molecular interactions through scientific publications. We hope that in the future such a model can be useful for automatically producing consensus views of molecular interaction data. Michael Krauthammer, Pauline Kra, Ivan Iossifov, Shawn M. Gomez, George Hripcsak, Vasileios Hatzivassiloglou, Carol Friedman, Andrey Rzhetsky |
ISMB | 1 |
| 2001 | A knowledge model for the interpretation and visualization of NLP-parsed discharged summaries
Michael Krauthammer, George Hripcsak |
AMIA | 1 |
| 2001 | Linking Protein Interaction Data to the MESH Hierarchy
Michael Krauthammer, Pauline Kra, Carol Friedman |
AMIA | 1 |
| 2000 | Using BLAST, A DNA and Protein Sequence Comparison Tool, for Finding Gene and Protein Names in Journal Articles
Michael Krauthammer, Andrey Rzhetsky, Pavel Morozov, Carol Friedman |
AMIA | 1 |
| 2000 | Knowledge Sharing in the Construction of an Anatomical Navigational Ontology
Michael Krauthammer, Nina Wacholder, Stephen B. Johnson, Gai Elhanan, Pat Molholt |
AMIA | 1 |
| 2000 | Designing a Navigational Ontology for Browsing and Accessing 3D Anatomical Images
Nina Wacholder, Judith M. Venuti, Michael Krauthammer, Pat Molholt |
AMIA | 3 |
| 2000 | A knowledge model for analysis and simulation of regulatory networksabstractMOTIVATION: In order to aid in hypothesis-driven experimental gene discovery, we are designing a computer application for the automatic retrieval of signal transduction data from electronic versions of scientific publications using natural language processing (NLP) techniques, as well as for visualizing and editing representations of regulatory systems. These systems describe both signal transduction and biochemical pathways within complex multicellular organisms, yeast, and bacteria. This computer application in turn requires the development of a domain-specific ontology, or knowledge model. RESULTS: We introduce an ontological model for the representation of biological knowledge related to regulatory networks in vertebrates. We outline a taxonomy of the concepts, define their 'whole-to-part' relationships, describe the properties of major concepts, and outline a set of the most important axioms. The ontology is partially realized in a computer system designed to aid researchers in biology and medicine in visualizing and editing a representation of a signal transduction system. Andrey Rzhetsky, Tomohiro Koike, Sergey Kalachikov, Shawn M. Gomez, Michael Krauthammer, Sabina H. Kaplan, Pauline Kra, James J. Russo, Carol Friedman |
Bioinform. | 5 |