VLDB 2026 Research / reviewers in the wild / expert
Predrag Radivojac
dblp:93/308
· DBLP profile ↗
51ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0002-6769-0793ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 39 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 7 · 2 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Are Your Fairness Metrics Accurate? A Semi-Supervised Approach to Improving Fairness Estimates Under Sample Selection BiasabstractA key challenge impeding the widespread deployment of machine learning is overcoming the impact of statistical biases in the data. Models trained on unrepresentative data can perform worse than anticipated and differentially affect cross-sections of the population. Therefore, evaluating and vetting models based on an appropriate notion of fairness is often indispensable, making accurate estimation of fairness metrics a critical step to safeguard against deployment of unfair algorithms. It is often assumed that a fairness metric computed from the observed data is accurate. However, in presence of selection bias, also referred to as distributional shifts, fairness metric estimates too can have systematic application-specific errors. In this work we demonstrate this phenomenon and, relying on access to an unbiased unlabeled data, derive a semi-supervised approach to mitigate estimation errors emerging from the biased labeled data. Specifically, we introduce a novel selection bias model called ''sub-class-conditional invariance'' (SCC-invariance), that offers a flexible framework to effectively capture distributional shifts in the real-world data, particularly compared to traditional models such as label shift and covariate shift. Assuming a finite Gaussian mixture form for each class-conditional distribution, we then derive an Expectation-Maximization algorithm to estimate model parameters and correction weights necessary for computing unbiased estimates. We focus on three widely used fairness metrics--equal opportunity, predictive equality, and predictive parity--and demonstrate the effectiveness of our approach in improving their estimates on synthetic data. Finally, we apply our bias mitigation approach to clinical genetics and study the fairness of pathogenicity predictors across ancestral groups. M. Clara De Paolis Kaluza, Thulasi Tholeti, Yile Chen 0006, Ricardo Baeza-Yates, Predrag Radivojac, Shantanu Jain |
KDD (2) | 5 |
| 2025 | Calibration of variant effect predictors for in-frame indels for clinical variant classificationabstractAbstract Insertions and deletions (indels) are the second most common type of variant in humans and are associated with a wide range of functional consequences. Despite their biological importance, indels, particularly non-frameshifting ones, remain understudied. While many computational predictors for missense variants have been rigorously evaluated, the clinical utility of tools for non-frameshifting indels remains uncertain. In this work, we calibrate non- frameshifting indel predictors for clinical variant classification, using a previously established framework developed for calibrating missense predictors. We constructed a dataset of high- confidence in-frame indel variants from ClinVar and gnomAD and estimated the prior probability of pathogenicity for all in-frame indels, as well as for insertions and deletions separately. Using a statistical framework based on local posterior probabilities, we then established score thresholds for several computational tools, including MutPred-Indel, VEST-indel, and CADD, corresponding to different levels of evidence for pathogenicity and benignity according to ACMG/AMP guidelines. These guidelines outline criteria for the clinical classification of variants, where different types of evidence (functional, population, computational, etc.) are weighted by strength (supporting, moderate, strong, etc.) and assigned point values that are summed to determine whether a variant is pathogenic, benign, or of uncertain significance. Distinct from several missense predictors which can reach the strong (+4 points) level of evidence, we find that most in-frame indel predictors reach the moderate (+2 points) or supporting (+1 point) evidence levels, demonstrating their clinical value, while highlighting the need for improved tools for indel interpretation. Haneen Abderrazzaq, Mugdha Singh, Timothy Bergquist, Vikas Pejaver, Anne H. O'Donnell-Luria, Predrag Radivojac |
Briefings Bioinform. | 6 |
| 2024 | Modeling Multiple Adverse Pregnancy Outcomes: Learning from Diverse Data Sources
Saurabh Mathur 0002, Veerendra P. Gadekar, Rashika Ramola, Ramachandran Thiruvengadam, David M. Haas, Shinjini Bhatnagar, Nitya Wadhwa, Garbhini Study Group, Predrag Radivojac, Himanshu Sinha, Kristian Kersting, Sriraam Natarajan |
AIME (1) | 10 |
| 2024 | An algorithm for decoy-free false discovery rate estimation in XL-MS/MS proteomicsabstractMOTIVATION: Cross-linking tandem mass spectrometry (XL-MS/MS) is an established analytical platform used to determine distance constraints between residues within a protein or from physically interacting proteins, thus improving our understanding of protein structure and function. To aid biological discovery with XL-MS/MS, it is essential that pairs of chemically linked peptides be accurately identified, a process that requires: (i) database search, that creates a ranked list of candidate peptide pairs for each experimental spectrum and (ii) false discovery rate (FDR) estimation, that determines the probability of a false match in a group of top-ranked peptide pairs with scores above a given threshold. Currently, the only available FDR estimation mechanism in XL-MS/MS is the target-decoy approach (TDA). However, despite its simplicity, TDA has both theoretical and practical limitations that impact the estimation accuracy and increase run time over potential decoy-free approaches (DFAs). RESULTS: We introduce a novel decoy-free framework for FDR estimation in XL-MS/MS. Our approach relies on multi-sample mixtures of skew normal distributions, where the latent components correspond to the scores of correct peptide pairs (both peptides identified correctly), partially incorrect peptide pairs (one peptide identified correctly, the other incorrectly), and incorrect peptide pairs (both peptides identified incorrectly). To learn these components, we exploit the score distributions of first- and second-ranked peptide-spectrum matches for each experimental spectrum and subsequently estimate FDR using a novel expectation-maximization algorithm with constraints. We evaluate the method on ten datasets and provide evidence that the proposed DFA is theoretically sound and a viable alternative to TDA owing to its good performance in terms of accuracy, variance of estimation, and run time. AVAILABILITY AND IMPLEMENTATION: https://github.com/shawn-peng/xlms. Yisu Peng, Shantanu Jain, Predrag Radivojac |
Bioinform. | 3 |
| 2023 | Leveraging Structure for Improved Classification of Grouped Biased DataabstractWe consider semi-supervised binary classification for applications in which data points are naturally grouped (e.g., survey responses grouped by state) and the labeled data is biased (e.g., survey respondents are not representative of the population). The groups overlap in the feature space and consequently the input-output patterns are related across the groups. To model the inherent structure in such data, we assume the partition-projected class-conditional invariance across groups, defined in terms of the group-agnostic feature space. We demonstrate that under this assumption, the group carries additional information about the class, over the group-agnostic features, with provably improved area under the ROC curve. Further assuming invariance of partition-projected class-conditional distributions across both labeled and unlabeled data, we derive a semi-supervised algorithm that explicitly leverages the structure to learn an optimal, group-aware, probability-calibrated classifier, despite the bias in the labeled data. Experiments on synthetic and real data demonstrate the efficacy of our algorithm over suitable baselines and ablative models, spanning standard supervised and semi-supervised learning approaches, with and without incorporating the group directly as a feature. Daniel Zeiberg, Shantanu Jain, Predrag Radivojac |
AAAI | 3 |
| 2021 | A Probabilistic Approach to Extract Qualitative Knowledge for Early Prediction of Gestational Diabetes
Athresh Karanam, Alexander L. Hayes, Harsha Kokel, David M. Haas, Predrag Radivojac, Sriraam Natarajan |
AIME | 5 |
| 2021 | Classification in biological networks with hypergraphlet kernelsabstractMOTIVATION: Biological and cellular systems are often modeled as graphs in which vertices represent objects of interest (genes, proteins and drugs) and edges represent relational ties between these objects (binds-to, interacts-with and regulates). This approach has been highly successful owing to the theory, methodology and software that support analysis and learning on graphs. Graphs, however, suffer from information loss when modeling physical systems due to their inability to accurately represent multiobject relationships. Hypergraphs, a generalization of graphs, provide a framework to mitigate information loss and unify disparate graph-based methodologies. RESULTS: We present a hypergraph-based approach for modeling biological systems and formulate vertex classification, edge classification and link prediction problems on (hyper)graphs as instances of vertex classification on (extended, dual) hypergraphs. We then introduce a novel kernel method on vertex- and edge-labeled (colored) hypergraphs for analysis and learning. The method is based on exact and inexact (via hypergraph edit distances) enumeration of hypergraphlets; i.e. small hypergraphs rooted at a vertex of interest. We empirically evaluate this method on fifteen biological networks and show its potential use in a positive-unlabeled setting to estimate the interactome sizes in various species. AVAILABILITY AND IMPLEMENTATION: https://github.com/jlugomar/hypergraphlet-kernels. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jose Lugo-Martinez, Daniel Zeiberg, Thomas Gaudelet, Noël Malod-Dognin, Natasa Przulj, Predrag Radivojac |
Bioinform. | 6 |
| 2020 | Class Prior Estimation with Biased Positives and Unlabeled ExamplesabstractPositive-unlabeled learning is often studied under the assumption that the labeled positive sample is drawn randomly from the true distribution of positives. In many application domains, however, certain regions in the support of the positive class-conditional distribution are over-represented while others are under-represented in the positive sample. Although this introduces problems in all aspects of positive-unlabeled learning, we begin to address this challenge by focusing on the estimation of class priors, quantities central to the estimation of posterior probabilities and the recovery of true classification performance. We start by making a set of assumptions to model the sampling bias. We then extend the identifiability theory of class priors from the unbiased to the biased setting. Finally, we derive an algorithm for estimating the class priors that relies on clustering to decompose the original problem into subproblems of unbiased positive-unlabeled learning. Our empirical investigation suggests feasibility of the correction strategy and overall good performance. Shantanu Jain, Justin Delano, Predrag Radivojac |
AAAI | 4 |
| 2020 | Fast Nonparametric Estimation of Class Proportions in the Positive-Unlabeled Classification SettingabstractEstimating class proportions has emerged as an important direction in positive-unlabeled learning. Well-estimated class priors are key to accurate approximation of posterior distributions and are necessary for the recovery of true classification performance. While significant progress has been made in the past decade, there remains a need for accurate strategies that scale to big data. Motivated by this need, we propose an intuitive and fast nonparametric algorithm to estimate class proportions. Unlike any of the previous methods, our algorithm uses a sampling strategy to repeatedly (1) draw an example from the set of positives, (2) record the minimum distance to any of the unlabeled examples, and (3) remove the nearest unlabeled example. We show that the point of sharp increase in the recorded distances corresponds to the desired proportion of positives in the unlabeled set and train a deep neural network to identify that point. Our distance-based algorithm is evaluated on forty datasets and compared to all currently available methods. We provide evidence that this new approach results in the most accurate performance and can be readily used on large datasets. Daniel Zeiberg, Shantanu Jain, Predrag Radivojac |
AAAI | 3 |
| 2020 | An examination of citation-based impact of the computational biology conferences
Jayvardan S. Naidu, Justin Delano, Scott Mathews, Predrag Radivojac |
Bioinform. | 4 |
| 2020 | New mixture models for decoy-free false discovery rate estimation in mass spectrometry proteomicsabstractMOTIVATION: Accurate estimation of false discovery rate (FDR) of spectral identification is a central problem in mass spectrometry-based proteomics. Over the past two decades, target-decoy approaches (TDAs) and decoy-free approaches (DFAs) have been widely used to estimate FDR. TDAs use a database of decoy species to faithfully model score distributions of incorrect peptide-spectrum matches (PSMs). DFAs, on the other hand, fit two-component mixture models to learn the parameters of correct and incorrect PSM score distributions. While conceptually straightforward, both approaches lead to problems in practice, particularly in experiments that push instrumentation to the limit and generate low fragmentation-efficiency and low signal-to-noise-ratio spectra. RESULTS: We introduce a new decoy-free framework for FDR estimation that generalizes present DFAs while exploiting more search data in a manner similar to TDAs. Our approach relies on multi-component mixtures, in which score distributions corresponding to the correct PSMs, best incorrect PSMs and second-best incorrect PSMs are modeled by the skew normal family. We derive EM algorithms to estimate parameters of these distributions from the scores of best and second-best PSMs associated with each experimental spectrum. We evaluate our models on multiple proteomics datasets and a HeLa cell digest case study consisting of more than a million spectra in total. We provide evidence of improved performance over existing DFAs and improved stability and speed over TDAs without any performance degradation. We propose that the new strategy has the potential to extend beyond peptide identification and reduce the need for TDA on all analytical platforms. AVAILABILITYAND IMPLEMENTATION: https://github.com/shawn-peng/FDR-estimation. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yisu Peng, Shantanu Jain, Yong Fuga Li, Michal Gregus, Alexander R. Ivanov, Olga Vitek, Predrag Radivojac |
Bioinform. | 7 |
| 2020 | The ortholog conjecture revisited: the value of orthologs and paralogs in function predictionabstractMOTIVATION: The computational prediction of gene function is a key step in making full use of newly sequenced genomes. Function is generally predicted by transferring annotations from homologous genes or proteins for which experimental evidence exists. The 'ortholog conjecture' proposes that orthologous genes should be preferred when making such predictions, as they evolve functions more slowly than paralogous genes. Previous research has provided little support for the ortholog conjecture, though the incomplete nature of the data cast doubt on the conclusions. RESULTS: We use experimental annotations from over 40 000 proteins, drawn from over 80 000 publications, to revisit the ortholog conjecture in two pairs of species: (i) Homo sapiens and Mus musculus and (ii) Saccharomyces cerevisiae and Schizosaccharomyces pombe. By making a distinction between questions about the evolution of function versus questions about the prediction of function, we find strong evidence against the ortholog conjecture in the context of function prediction, though questions about the evolution of function remain difficult to address. In both pairs of species, we quantify the amount of information that would be ignored if paralogs are discarded, as well as the resulting loss in prediction accuracy. Taken as a whole, our results support the view that the types of homologs used for function transfer are largely irrelevant to the task of function prediction. Maximizing the amount of data used for this task, regardless of whether it comes from orthologs or paralogs, is most likely to lead to higher prediction accuracy. AVAILABILITY AND IMPLEMENTATION: https://github.com/predragradivojac/oc. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Moses Stamboulian, Rafael F. Guerrero, Matthew W. Hahn, Predrag Radivojac |
Bioinform. | 4 |
| 2019 | ISMB/ECCB 2019 ProceedingsabstractThe biennial joint meeting of ISMB (27th Annual Conference on Intelligent Systems for Molecular Biology) and ECCB (18th European Conference on Computational Biology) was held in Basel, Switzerland, July 21–25, 2019. ISMB is the flagship conference of the International Society for Computational Biology and the world’s premier forum for dissemination of scientific research in computational biology and its intersection with other areas. ECCB is similarly a top venue in the field, with a long tradition of publishing and presenting world-class research. This special issue serves as the Proceedings of ISMB/ECCB 2019. Following a successful model with a centralized manuscript review and acceptance process, this year’s conference organization provided the community with a unified submission interface for high-quality papers in the field of computational biology. The review process across 10 scientific areas was supervised by the Senior Program Committee (SPC), consisting of the Proceedings Chairs and Area Chairs (AC). About a third of the ACs were nominated by the Communities of Special Interests (COSIs), reflecting the desire of the ISMB/ECCB 2019 Steering Committee to involve COSIs in conference organization and the review process. Overall, the SPC consisted of 21 individuals; see Table 1. Thematic areas of ISMB/ECCB 2019 Note: The table lists the ACs for each theme, the number of reviewed papers, the number of accepted papers, and the acceptance rate for each area. A special area of General Computational Biology was created for papers not fitting in any of the predefined areas. Thematic areas of ISMB/ECCB 2019 Note: The table lists the ACs for each theme, the number of reviewed papers, the number of accepted papers, and the acceptance rate for each area. A special area of General Computational Biology was created for papers not fitting in any of the predefined areas. The scope of the conference includes theoretical papers, algorithms and statistical methods that allow for important novel biological insights and broadly defined intellectual contributions. We invited submission of papers in nine general scientific areas, organized by the relevant biological problems (Table 1). The 10th area of General Computational Biology was created to accommodate innovation outside of specified fields. The papers submitted to this area were largely in Text Mining, Mass Spectrometry and Visualization subfields, indicating community interest in these areas. All papers submitted to all areas were expected to present methodological and scientific contributions to the specific areas of submission. Abstracts, previously published papers, position papers, perspectives and reviews are not eligible for submission to the ISMB/ECCB Proceedings track. In total, 366 papers were submitted—a 10.6% increase over ISMB 2018. Of these, 363 papers were sent to review, receiving 1303 reviews from 388 Program Committee members. This constitutes an average of 3.6 reviews per submission. Three submissions had two completed reviews, 175 had three, 157 had four, 24 had five and four submissions had six completed reviews. The conditional acceptance information was provided to the authors a month after the submission deadline and final acceptance conveyed in another month. Overall, 69 manuscripts were accepted for a final acceptance rate of 18.9%. The distribution of papers reviewed in different areas is shown in Table 1. As was the case in 2018, the authors had the opportunity to request that their accepted manuscripts be presented in one of the COSI sessions. The most requests were made for MLCSB (64), followed by NetBio (39), Evolution (37), RegSys (37), TransMed (34), HiTSeq (33), Function (30), Microbiome (18), VarI (16), RNA (15), Text Mining (11), BioVis (10), Bio-Ontologies (7), CAMDA (4), Education (2) and CompMS (2). A further 35% of the submitted manuscripts made no specific COSI session request. Final assignments to COSI sessions were made by the Proceedings Chairs on the basis of the author and COSI requests. Three papers were assigned to the General Computational Biology session for presentation. The distribution of accepted papers to COSIs is shown in Table 2. Distribution of ISMB/ECCB 2019 Proceedings papers to COSIs Distribution of ISMB/ECCB 2019 Proceedings papers to COSIs We thank the ISMB/ECCB 2019 Steering Committee for their support and guidance. Special thanks go to the SPC and reviewers for their fantastic work dedicated to maintaining the high-quality standards of ISMB/ECCB in a compressed time frame. We are particularly grateful to Steven Leard, Diane Kovats and Pat Rodenburg for world-class organizational support. We would also like to thank the COSIs for nominating the ACs and the COSI contacts for their help in identifying reviewers and incorporating accepted papers into their programs. Finally, we thank the community for their interest and engagement in this conference—ISMB/ECCB 2019 belongs to you! With this, we invite you to read the Proceedings of ISMB/ECCB 2019. See you next year in Montreal, Canada, for ISMB 2020. Conflict of Interest: none declared. Yana Bromberg, Nadia El-Mabrouk, Predrag Radivojac |
Bioinform. | 3 |
| 2019 | A new class of metrics for learning on real-valued and structured data
Ruiyu Yang, Yuxiang Jiang, Scott Mathews, Elizabeth A. Housworth, Matthew W. Hahn, Predrag Radivojac |
Data Min. Knowl. Discov. | 6 |
| 2019 | Pathogenicity and functional impact of non-frameshifting insertion/deletion variation in the human genomeabstractDifferentiation between phenotypically neutral and disease-causing genetic variation remains an open and relevant problem. Among different types of variation, non-frameshifting insertions and deletions (indels) represent an understudied group with widespread phenotypic consequences. To address this challenge, we present a machine learning method, MutPred-Indel, that predicts pathogenicity and identifies types of functional residues impacted by non-frameshifting insertion/deletion variation. The model shows good predictive performance as well as the ability to identify impacted structural and functional residues including secondary structure, intrinsic disorder, metal and macromolecular binding, post-translational modifications, allosteric sites, and catalytic residues. We identify structural and functional mechanisms impacted preferentially by germline variation from the Human Gene Mutation Database, recurrent somatic variation from COSMIC in the context of different cancers, as well as de novo variants from families with autism spectrum disorder. Further, the distributions of pathogenicity prediction scores generated by MutPred-Indel are shown to differentiate highly recurrent from non-recurrent somatic variation. Collectively, we present a framework to facilitate the interrogation of both pathogenicity and the functional effects of non-frameshifting insertion/deletion variants. The MutPred-Indel webserver is available at http://mutpred.mutdb.org/. Kymberleigh A. Pagel, Danny Antaki, Aojie Lian, Matthew E. Mort, David N. Cooper, Jonathan Sebat, Lilia M. Iakoucheva, Sean D. Mooney, Predrag Radivojac |
PLoS Comput. Biol. | 9 |
| 2018 | On Whom Should I Perform this Lab Test Next? An Active Feature Elicitation ApproachabstractWe consider the problem of actively feature elicitation in which given a few examples with all the features (say the full EHR) and a few examples with some of the features (say demographics), the goal is to identify the set of examples on whom more information (say the lab tests) needs to be collected. The observation is that some set of features may be more expensive, personal or cumbersome to collect. We propose an active learning approach which identifies examples that are dissimilar to the ones with the full set of data and acquire the complete set of features for these examples. Motivated by real clinical tasks, our extensive evaluation on three clinical tasks demonstrate the effectiveness of this approach. Sriraam Natarajan, Srijita Das 0001, Nandini Ramanan, Gautam Kunapuli, Predrag Radivojac |
IJCAI | 5 |
| 2018 | ISMB 2018 proceedingsabstractThe 26th Annual Conference on Intelligent Systems for Molecular Biology (ISMB) was held in Chicago, Illinois, USA, July 6–10, 2018. ISMB is the flagship conference of the International Society for Computational Biology (ISCB) and the world’s premier forum for dissemination of scientific research in computational biology and its intersection with other fields. This special issue serves as the Proceedings of ISMB 2018. One of the primary goals of this year’s conference organization was to provide the community with a unified interface for submission/acceptance of high-quality papers in the field of computational biology. That is, we explicitly aimed to ensure homogeneity in quality of accepted papers across different scientific themes. Moreover, we looked to simultaneously acquaint researchers outside the immediate ISCB community with the areas of interest represented by our Community of Special Interests (COSIs; COSI stands for the Community of Special Interests. There are 20 active COSIs currently operating under ISCB. Of those, 16 COSIs participated in the ISMB 2018 Proceedings.) and to invite submissions beyond these pre-determined interests. The scope of ISMB includes theoretical papers, algorithms, or statistical methods, while allowing for important novel biological insights and broadly defined intellectual contributions. We invited submission of papers in 10 general scientific areas (Table 1), organized by relevant biological problems. The area of General Computational Biology was created to accommodate under-represented innovation in the field. Upon inspection of the papers submitted to this area, three methodology-specific scientific sub-areas were identified (Text Mining, Mass Spectrometry and Visualization) in addition to a number of unclassifiable outliers. All papers submitted to all areas were expected to present methodological and scientific contributions to the specific areas of submission. Thematic areas of ISMB 2018 Notes: The table lists the ACs for each theme, the number of reviewed papers, the number of accepted papers and the acceptance rate for each area. A special area of General Computational Biology was created for papers not fitting in any of the pre-defined areas. This area contained 6 Text Mining papers (1 accepted), 7 Mass Spectrometry papers (2 accepted), 5 Visualization papers (2 accepted) and 14 additional diverse papers (1 accepted). The list of CCs includes: Rolf Backhofen, Jan Baumbach, Emidio Capriotti, Ana Conesa, Christophe Dessimoz, Jeremy Goecks, Pavel Labaj, Florian Markowetz, Nicola Mulder, Alice McHardy, Natasa Przulj, Venkata Satagopam, Saurabh Sinha, Karin Verspoor and Olga Vitek. Thematic areas of ISMB 2018 Notes: The table lists the ACs for each theme, the number of reviewed papers, the number of accepted papers and the acceptance rate for each area. A special area of General Computational Biology was created for papers not fitting in any of the pre-defined areas. This area contained 6 Text Mining papers (1 accepted), 7 Mass Spectrometry papers (2 accepted), 5 Visualization papers (2 accepted) and 14 additional diverse papers (1 accepted). The list of CCs includes: Rolf Backhofen, Jan Baumbach, Emidio Capriotti, Ana Conesa, Christophe Dessimoz, Jeremy Goecks, Pavel Labaj, Florian Markowetz, Nicola Mulder, Alice McHardy, Natasa Przulj, Venkata Satagopam, Saurabh Sinha, Karin Verspoor and Olga Vitek. Prominent researchers were selected to the Senior Program Committee (SPC). SPC consisted of the Proceedings Chairs, Area Chairs (ACs) and COSI-selected Chairs (CCs). This structure reflected the desire of the ISMB 2018 Steering Committee to involve COSIs early on and throughout the review process, in part because the Proceedings papers were to be ultimately presented in COSI sessions at the main meeting. Overall, the SPC consisted of 26 people (Table 1). In total, 331 papers were submitted for an 18.6% increase over ISMB/ECCB 2017. Of those, 324 papers were sent to review for which 1040 reviews were written by 337 Program Committee (PC) members (average of 3.2 reviews per submission). Twenty-nine submissions had 2 completed reviews, 209 had 3, 64 had 4, 15 had 5 and 4 submissions had 6 completed reviews. The conditional acceptance information was provided to the authors a month later and final acceptance conveyed 2 months after initial submission. Overall, 65 manuscripts were accepted for a final acceptance rate of 19.6%. The distribution of papers reviewed in different areas is shown in Table 1. A new feature of this year’s ISMB submission was that the authors had an opportunity to request that their accepted manuscript be presented in one of the COSI sessions. The most requests were made for MLCSB (73), followed by NetBio (47), HiTSeq (29), Function (29), RegSys (27), TransMed (25), RNA (23), Evolution (20), Bio-Ontologies (17), Microbiome (15), BioVis (13), VarI (5), CAMDA (5) and Education (1). Roughly 37% of the submitted manuscripts did not provide this information. Final assignments to COSI sessions were made by the Proceedings Chairs on the basis of the author and COSI requests. One of the papers did not fit any of the COSIs and was assigned to the General Computational Biology session. The distribution of accepted papers to COSIs is shown in Table 2. Distribution of ISMB 2018 proceedings papers to COSIs Distribution of ISMB 2018 proceedings papers to COSIs We thank the ISMB 2018 Steering Committee for their guidance. Special thanks go to the SPC and reviewers for all their fantastic work dedicated to maintaining the high quality standards of ISMB in a very short timeframe. We are particularly grateful to Steven Leard, Diane Kovats and Pat Rodenburg for world-class organizational support. Finally, we thank the community for their interest and engagement in this conference––ISMB 2018 belongs to you! With this, we invite you to read the Proceedings of ISMB 2018. See you next year in Basel, Switzerland, for ISMB/ECCB 2019. Conflict of Interest: none declared. Yana Bromberg, Predrag Radivojac |
Bioinform. | 2 |
| 2018 | Enumerating consistent sub-graphs of directed acyclic graphs: an insight into biomedical ontologiesabstractMotivation: Modern problems of concept annotation associate an object of interest (gene, individual, text document) with a set of interrelated textual descriptors (functions, diseases, topics), often organized in concept hierarchies or ontologies. Most ontology can be seen as directed acyclic graphs (DAGs), where nodes represent concepts and edges represent relational ties between these concepts. Given an ontology graph, each object can only be annotated by a consistent sub-graph; that is, a sub-graph such that if an object is annotated by a particular concept, it must also be annotated by all other concepts that generalize it. Ontologies therefore provide a compact representation of a large space of possible consistent sub-graphs; however, until now we have not been aware of a practical algorithm that can enumerate such annotation spaces for a given ontology. Results: We propose an algorithm for enumerating consistent sub-graphs of DAGs. The algorithm recursively partitions the graph into strictly smaller graphs until the resulting graph becomes a rooted tree (forest), for which a linear-time solution is computed. It then combines the tallies from graphs created in the recursion to obtain the final count. We prove the correctness of this algorithm, propose several practical accelerations, evaluate it on random graphs and then apply it to characterize four major biomedical ontologies. We believe this work provides valuable insights into the complexity of concept annotation spaces and its potential influence on the predictability of ontological annotation. Availability and implementation: https://github.com/shawn-peng/counting-consistent-sub-DAG. Supplementary information: Supplementary data are available at Bioinformatics online. Yisu Peng, Yuxiang Jiang, Predrag Radivojac |
Bioinform. | 3 |
| 2018 | Proteomic Evidence for In-Frame and Out-of-Frame Alternatively Spliced Isoforms in Human and MouseabstractIn order to find evidence for translation of alternatively spliced transcripts, especially those that result in a change in reading frame, we collected exon-skipping cases previously found by RNA-Seq and applied a computational approach to screen millions of mass spectra. These spectra came from seven human and six mouse tissues, five of which are the same between the two organisms: liver, kidney, lung, heart, and brain. Overall, we detected 4 percent of all exon-skipping events found in RNA-seq data, regardless of their effect on reading frame. The fraction of alternative isoforms detected did not differ between out-of-frame and in-frame events. Moreover, the fraction of identified alternative exon-exon junctions and constitutive junctions were similar. Together, our results suggest that both in-frame and out-of-frame translation may be actively used to regulate protein activity or localization. Rodrigo F. Ramalho, Sujun Li, Predrag Radivojac, Matthew W. Hahn |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2018 | Ultra High-Dimensional Nonlinear Feature Selection for Big Biological DataabstractMachine learning methods are used to discover complex nonlinear relationships in biological and medical data. However, sophisticated learning models are computationally unfeasible for data with millions of features. Here, we introduce the first feature selection method for nonlinear learning problems that can scale up to large, ultra-high dimensional biological data. More specifically, we scale up the novel Hilbert-Schmidt Independence Criterion Lasso (HSIC Lasso) to handle millions of features with tens of thousand samples. The proposed method is guaranteed to find an optimal subset of maximally predictive features with minimal redundancy, yielding higher predictive power and improved interpretability. Its effectiveness is demonstrated through applications to classify phenotypes based on module expression in human prostate cancer patients and to detect enzymes among protein structures. We achieve high accuracy with as few as 20 out of one million features-a dimensionality reduction of 99.998 percent. Our algorithm can be implemented on commodity cloud computing platforms. The dramatic reduction of features may lead to the ubiquitous deployment of sophisticated prediction models in mobile health care applications. Makoto Yamada, Jiliang Tang, Jose Lugo-Martinez, Ermin Hodzic, Raunak Shrestha, Avishek Saha, Hua Ouyang, Dawei Yin 0001, Hiroshi Mamitsuka, Süleyman Cenk Sahinalp, Predrag Radivojac, Filippo Menczer, Yi Chang 0001 |
IEEE Trans. Knowl. Data Eng. | 11 |
| 2017 | Recovering True Classifier Performance in Positive-Unlabeled LearningabstractA common approach in positive-unlabeled learning is to train a classification model between labeled and unlabeled data. This strategy is in fact known to give an optimal classifier under mild conditions; however, it results in biased empirical estimates of the classifier performance. In this work, we show that the typically used performance measures such as the receiver operating characteristic curve, or the precision recall curve obtained on such data can be corrected with the knowledge of class priors; i.e., the proportions of the positive and negative examples in the unlabeled data. We extend the results to a noisy setting where some of the examples labeled positive are in fact negative and show that the correction also requires the knowledge of the proportion of noisy examples in the labeled positives. Using state-of-the-art algorithms to estimate the positive class prior and the proportion of noise, we experimentally evaluate two correction approaches and demonstrate their efficacy on real-life data. Shantanu Jain, Martha White, Predrag Radivojac |
AAAI | 3 |
| 2017 | When loss-of-function is loss of function: assessing mutational signatures and impact of loss-of-function genetic variantsabstractMOTIVATION: Loss-of-function genetic variants are frequently associated with severe clinical phenotypes, yet many are present in the genomes of healthy individuals. The available methods to assess the impact of these variants rely primarily upon evolutionary conservation with little to no consideration of the structural and functional implications for the protein. They further do not provide information to the user regarding specific molecular alterations potentially causative of disease. RESULTS: To address this, we investigate protein features underlying loss-of-function genetic variation and develop a machine learning method, MutPred-LOF, for the discrimination of pathogenic and tolerated variants that can also generate hypotheses on specific molecular events disrupted by the variant. We investigate a large set of human variants derived from the Human Gene Mutation Database, ClinVar and the Exome Aggregation Consortium. Our prediction method shows an area under the Receiver Operating Characteristic curve of 0.85 for all loss-of-function variants and 0.75 for proteins in which both pathogenic and neutral variants have been observed. We applied MutPred-LOF to a set of 1142 de novo vari3ants from neurodevelopmental disorders and find enrichment of pathogenic variants in affected individuals. Overall, our results highlight the potential of computational tools to elucidate causal mechanisms underlying loss of protein function in loss-of-function variants. AVAILABILITY AND IMPLEMENTATION: http://mutpred.mutdb.org. CONTACT: [email protected]. Kymberleigh A. Pagel, Vikas Pejaver, Guan Ning Lin, Hyun-Jun Nam, Matthew E. Mort, David N. Cooper, Jonathan Sebat, Lilia M. Iakoucheva, Sean D. Mooney, Predrag Radivojac |
Bioinform. | 10 |
| 2016 | Estimating the class prior and posterior from noisy positives and unlabeled dataabstractWe develop a classification algorithm for estimating posterior distributions from positive-unlabeled data, that is robust to noise in the positive labels and effective for high-dimensional data. In recent years, several algorithms have been proposed to learn from positive-unlabeled data; however, many of these contributions remain theoretical, performing poorly on real high-dimensional data that is typically contaminated with noise. We build on this previous work to develop two practical classification algorithms that explicitly model the noise in the positive labels and utilize univariate transforms built on discriminative classifiers. We prove that these univariate transforms preserve the class prior, enabling estimation in the univariate space and avoiding kernel density estimation for high-dimensional data. The theoretical development and parametric and nonparametric algorithms proposed here constitute an important step towards wide-spread use of robust classification algorithms for positive-unlabeled data. Shantanu Jain, Martha White, Predrag Radivojac |
NIPS | 3 |
| 2016 | The Loss and Gain of Functional Amino Acid Residues Is a Common Mechanism Causing Human Inherited DiseaseabstractElucidating the precise molecular events altered by disease-causing genetic variants represents a major challenge in translational bioinformatics. To this end, many studies have investigated the structural and functional impact of amino acid substitutions. Most of these studies were however limited in scope to either individual molecular functions or were concerned with functional effects (e.g. deleterious vs. neutral) without specifically considering possible molecular alterations. The recent growth of structural, molecular and genetic data presents an opportunity for more comprehensive studies to consider the structural environment of a residue of interest, to hypothesize specific molecular effects of sequence variants and to statistically associate these effects with genetic disease. In this study, we analyzed data sets of disease-causing and putatively neutral human variants mapped to protein 3D structures as part of a systematic study of the loss and gain of various types of functional attribute potentially underlying pathogenic molecular alterations. We first propose a formal model to assess probabilistically function-impacting variants. We then develop an array of structure-based functional residue predictors, evaluate their performance, and use them to quantify the impact of disease-causing amino acid substitutions on catalytic activity, metal binding, macromolecular binding, ligand binding, allosteric regulation and post-translational modifications. We show that our methodology generates actionable biological hypotheses for up to 41% of disease-causing genetic variants mapped to protein structures suggesting that it can be reliably used to guide experimental validation. Our results suggest that a significant fraction of disease-causing human variants mapping to protein structures are function-altering both in the presence and absence of stability disruption. Jose Lugo-Martinez, Vikas Pejaver, Kymberleigh A. Pagel, Shantanu Jain, Matthew E. Mort, David N. Cooper, Sean D. Mooney, Predrag Radivojac |
PLoS Comput. Biol. | 8 |
| 2016 | Applying, Evaluating and Refining Bioinformatics Core Competencies (An Update from the Curriculum Task Force of ISCB's Education Committee)abstractThe Curriculum Task Force (CTF) of ISCB’s Education Committee seeks to define curricular guidelines for those who educate or train bioinformatics professionals at all career stages. A recent report of the CTF [1] presented a draft set of bioinformatics core competencies, derived from the results of surveys of (1) core facility directors, (2) career opportunities, and (3) existing curricula.
Since the publication of its 2014 report, the CTF has focused on the application of the guidelines in varied contexts to identify areas where refinement is needed. As a first step, the task force held an open meeting at the ISMB conference in July 2014. The ideas discussed at the meeting spawned four working groups (WGs), which focus on (i) defining core competencies for specific types and levels of bioinformatics training, (ii) mapping the curriculum guidelines and competencies to existing materials in order to identify the need for development of new materials, and (iii) identifying where revision of the guidelines may be valuable. The CTF is engaging the ISCB community through open WG meetings at ISCB’s official conferences. Thus far, the WGs have convened at the ISCB Great Lakes Bioinformatics Conference (Purdue University, May 2015) and at the ISMB/ECCB Conference (Dublin, Ireland, July 2015). Additionally, the CTF held a workshop at the Annual General Meeting of the Global Organization of Bioinformatics Learning, Education and Training (Cape Town, South Africa, November 2015). Specifically, the draft competencies have been employed in a wide range of activities and contexts (see Table 1 and [2–11]), including the development of new curricula, the analysis of existing curricula, and the creation of new roles involving bioinformatics. These activities have resulted in the identification of several areas where refinement would be useful:
Table 1
Summary of the activities of the ISCB Curriculum Task Force.
Identify different levels or phases of competency. It would be helpful to define different phases of competency development, or different levels of competency appropriate for distinct roles.
Define competency profiles for disciplines that don’t fit into our current silos. Bioengineering provides an illustrative example of a discipline that requires core competency in bioinformatics but does not fit into our current categories. There are almost certainly others. It would be helpful if we could provide some guidance on how to produce ‘hybrid’ competency profiles, perhaps borrowing some competencies from the TF’s core set and others from different disciplines. The LifeTrain initiative (www.lifetrain.eu) [2, 3] is collecting competency profiles for a range of disciplines of relevance to the biomedical sciences and may provide a useful resource kit for this.
Broaden the scope of the competency profiles in response to cutting-edge and emerging research. Current areas requiring improvement include incorporating competencies that capture a fundamental understanding of the biological principles central to analyzing biomolecular data, and broadening the user WG to include applications beyond medicine.
Provide guidance on the evidence required to assess whether someone has acquired each competency. For undergraduate, Master’s and PhD programs, learning outcomes for each competency, perhaps with examples of appropriate means of assessment, would be valuable. For established professionals who need to assimilate competencies into their working lives, a different approach may be required (such as keeping a portfolio to capture evidence of competency); the CTF should seek guidance from relevant professional bodies, especially in regulated professions such as healthcare.
Provide indicative course content or examples of programs that map to the competency requirements. We do not wish to prescribe what course providers should teach or how they should teach it; however, if a course provider is designing a course to meet a specific competency requirement, it may be helpful to find examples of other programs that do this successfully. One way of achieving this is by mapping existing training content to the TF’s competencies. Another way might be to provide an indication, perhaps based on several courses, of the course content that would meet the competency requirements. This would give course providers the freedom to build their own course syllabi without having to reinvent the wheel. Initiatives to collect examples of Creative Commons (or otherwise reusable) course materials will provide an extremely valuable bank of training materials that could be mapped to the core competencies. Lonnie R. Welch, Catherine Brooksbank, Russell Schwartz, Sarah L. Morgan, Bruno A. Gaëta, Alastair M. Kilpatrick, Daniel Mietchen, Benjamin L. Moore, Nicola J. Mulder, Mark A. Pauley, William R. Pearson, Predrag Radivojac, Naomi Rosenberg, Anne G. Rosenwald, Gabriella Rustici, Tandy J. Warnow |
PLoS Comput. Biol. | 12 |
| 2015 | Ten Simple Rules for a Community Computational ChallengeabstractIn science, the relationship between methods and discovery is symbiotic.As we discover more, we are able to construct more precise and sensitive tools and methods that enable further discovery.With better lens crafting came microscopes, and with them the discovery of living cells.In the last 40 years, advances in molecular biology, statistics, and computer science have ushered in the field of bioinformatics and the genomic era.Computational scientists enjoy developing new methods, and the community encourages them to do so.Indeed, the editorial guidelines for PLOS Computational Biology require manuscripts to apply novel methods.However, it is often confusing to know which method to choose: which method is best?And, in this context, what does "best" mean?To help choose an appropriate method for a particular task, scientists often form community-based challenges for the unbiased evaluation of methods in a given field.These challenges help evaluate existing and novel methods, while helping to coalesce a community and leading to new ideas and collaborations.In computational biology, the first of these challenges was arguably the Critical Assessment of protein Structure Prediction, or CASP [1], whose goal is to evaluate methods for predicting three-dimensional protein structure from amino acid sequence.The first CASP meeting was held in December of 1994, following a "prediction period" where members of the community were presented with protein amino acid sequences and asked to predict their three dimensional structures.The sequences that were chosen had recently been solved by X-ray crystallography but had not been not published or released until after the predictions from the community were made.Since the first CASP, we have seen many successful challenges, including Critical Assessment of Function Annotation (CAFA) for protein function prediction [2], Critical Assessment of Genome Interpretation (CAGI) (for genome interpretation) [3], Critical Assessment of Massive (originally "Microarray") Data Analysis (CAMDA) (for large-scale biological data) [4], BioCreative (for biomedical text mining) [5], the Assemblathon (for sequence assembly), and the NCI-DREAM Challenges (for various biomedical challenges), amongst others [6].Computational challenges also help solve new problems.While the original CASP experiment was developed to evaluate existing methods applied to current problems, other communities often look at other areas for which there are no existing tools.These challenges have spread successfully to industry, and companies such as Innocentive [7] and X-Prize [8] offer large prizes for solving novel questions. Iddo Friedberg, Mark N. Wass, Sean D. Mooney, Predrag Radivojac |
PLoS Comput. Biol. | 4 |
| 2015 | A link graph-based approach to identify forum spamabstractABSTRACT Web spammers have taken note of the popularity of public forums such as blogs, wikis, webboards, and guestbooks. They are now exploiting them with the purpose of driving traffic to their malicious or fraudulent websites, such as those used for phishing, distributing malware, or selling counterfeit pharmaceuticals. A popular technique they use is to spam these forums with URLs to their spam websites. We consider the problem of classifying URLs posted to forums as spam or legitimate by considering the link structure of the graph rooted at the posted URL. We investigate various graph metrics and associated metadata to analyze link structures. To lessen noisy structural characteristics of the link graphs for spam classification, we also examine two techniques: differing depths and aggregating sub‐graphs of the link graphs. Our results show that a support vector machine classifier based on combinations of graph metrics and metadata of link graphs can achieve a pragmatically high performance in forum spam detection. Copyright © 2014 John Wiley & Sons, Ltd. Youngsang Shin, Steven Myers, Minaxi Gupta 0001, Predrag Radivojac |
Secur. Commun. Networks | 4 |
| 2014 | The impact of incomplete knowledge on the evaluation of protein function prediction: a structured-output learning perspectiveabstractMOTIVATION: The automated functional annotation of biological macromolecules is a problem of computational assignment of biological concepts or ontological terms to genes and gene products. A number of methods have been developed to computationally annotate genes using standardized nomenclature such as Gene Ontology (GO). However, questions remain about the possibility for development of accurate methods that can integrate disparate molecular data as well as about an unbiased evaluation of these methods. One important concern is that experimental annotations of proteins are incomplete. This raises questions as to whether and to what degree currently available data can be reliably used to train computational models and estimate their performance accuracy. RESULTS: We study the effect of incomplete experimental annotations on the reliability of performance evaluation in protein function prediction. Using the structured-output learning framework, we provide theoretical analyses and carry out simulations to characterize the effect of growing experimental annotations on the correctness and stability of performance estimates corresponding to different types of methods. We then analyze real biological data by simulating the prediction, evaluation and subsequent re-evaluation (after additional experimental annotations become available) of GO term predictions. Our results agree with previous observations that incomplete and accumulating experimental annotations have the potential to significantly impact accuracy assessments. We find that their influence reflects a complex interplay between the prediction algorithm, performance metric and underlying ontology. However, using the available experimental data and under realistic assumptions, our results also suggest that current large-scale evaluations are meaningful and almost surprisingly reliable. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yuxiang Jiang, Wyatt Travis Clark, Iddo Friedberg, Predrag Radivojac |
Bioinform. | 4 |
| 2014 | The automated function prediction SIG looks back at 2013 and prepares for 2014abstractAbstract Contact: [email protected] or [email protected] Mark N. Wass, Sean D. Mooney, Michal Linial, Predrag Radivojac, Iddo Friedberg |
Bioinform. | 4 |
| 2014 | Bioinformatics Curriculum Guidelines: Toward a Definition of Core CompetenciesabstractRapid advances in the life sciences and in related information technologies necessitate the ongoing refinement of bioinformatics educational programs in order to maintain their relevance. As the discipline of bioinformatics and computational biology expands and matures, it is important to characterize the elements that contribute to the success of professionals in this field. These individuals work in a wide variety of settings, including bioinformatics core facilities, biological and medical research laboratories, software development organizations, pharmaceutical and instrument development companies, and institutions that provide education, service, and training. In response to this need, the Curriculum Task Force of the International Society for Computational Biology (ISCB) Education Committee seeks to define curricular guidelines for those who train and educate bioinformaticians. The previous report of the task force summarized a survey that was conducted to gather input regarding the skill set needed by bioinformaticians [1]. The current article details a subsequent effort, wherein the task force broadened its perspectives by examining bioinformatics career opportunities, surveying directors of bioinformatics core facilities, and reviewing bioinformatics education programs. Lonnie R. Welch, Fran Lewitter, Russell Schwartz, Catherine Brooksbank, Predrag Radivojac, Bruno A. Gaëta, Maria Victoria Schneider |
PLoS Comput. Biol. | 5 |
| 2013 | Information-theoretic evaluation of predicted ontological annotationsabstractMOTIVATION: The development of effective methods for the prediction of ontological annotations is an important goal in computational biology, with protein function prediction and disease gene prioritization gaining wide recognition. Although various algorithms have been proposed for these tasks, evaluating their performance is difficult owing to problems caused both by the structure of biomedical ontologies and biased or incomplete experimental annotations of genes and gene products. RESULTS: We propose an information-theoretic framework to evaluate the performance of computational protein function prediction. We use a Bayesian network, structured according to the underlying ontology, to model the prior probability of a protein's function. We then define two concepts, misinformation and remaining uncertainty, that can be seen as information-theoretic analogs of precision and recall. Finally, we propose a single statistic, referred to as semantic distance, that can be used to rank classification models. We evaluate our approach by analyzing the performance of three protein function predictors of Gene Ontology terms and provide evidence that it addresses several weaknesses of currently used metrics. We believe this framework provides useful insights into the performance of protein function prediction tools. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Wyatt Travis Clark, Predrag Radivojac |
Bioinform. | 2 |
| 2013 | ISCB Computational Biology Wikipedia CompetitionabstractThe International Society for Computational Biology is pleased to announce the 2013 ISCB Computational Biology Wikipedia competition. The competition, in which entrants create or improve the content of any Wikipedia article in the field of computational biology, is open to all students and trainees. Further information about the competition can be found here: http://en.wikipedia.org/wiki/Wikipedia:WikiProject_Computational_Biology/ISCB_competition_announcement_2013
The mission of the ISCB is to promote the use of computational biology and to help educate the next generation of computational biologists. The society has numerous activities that help to address these aims, including conferences, training and mentoring initiatives, and an active student council.
As the world's largest online encyclopedia, Wikipedia has become an indispensable resource for those seeking information on all scientific and technical topics. The English language version of Wikipedia contains over 4.2 million articles, and Wikipedia is now available in 286 languages. The global rise in smartphone use, which allows access to Wikipedia, means that a large fraction of the world's population can now gain access to the world's knowledge. Wikipedia is the most successful example of crowd-sourcing with about 80,000 active editors updating its content each month.
But is Wikipedia a good source of information for computational biology? Certainly, many people are reading the articles. For example, the Bioinformatics article has been visited 1,600 times per day over the last 3 months. Wikipedia contains articles on algorithms, biological databases, software packages, and biographies of eminent computational biologists. The computational biology content ranges from incomplete, a mere “stub” of an article in Wikipedia parlance, to highly detailed Featured Articles. A group of Wikipedia editors have formed the Computational Biology Wikiproject (http://en.wikipedia.org/wiki/Wikipedia:WikiProject_Computational_Biology). This group oversees the computational biology articles and rates them for their importance and their quality. Figure 1 shows the current state of the articles (see also Figure S1). In total, there are over 1,140 articles that have been considered as falling under Computational Biology. There are a small number of articles that have been brought up to the highest levels of quality (Featured Article and Good Article) such as Multiple Sequence Alignment, Genome Wide Association Study, and Folding@home.
Figure 1
The computational biology articles rated by quality and importance by the Wikipedia Computational Biology Wikiproject.
The 2012 competition began 9th September 2012 (coinciding with the start of the European Conference on Computational Biology) and finished four months later on the 10th January 2013. Each article entered in the competition was reviewed for a difference in article quality between these two dates. In 2012, there were 13 substantive entries into the competition. Six of these articles were shortlisted by members of the ISCB Student Council and then considered by the judging panel. The judging panel considered articles based on the criteria of clarity of the writing, depth of knowledge of the subject, and quality of figures and images used. In one case, it was clear that the article was largely derived from a published review, and was not considered further. For the other entries, the quantity and quality of the contributions were very good, and it was a challenge to rank the articles. After much deliberation, the judging panel selected the following as the winners of the 2012 ISCB Wikipedia competition:
1st prize: James Estevez for improvements to the Genomics Article.
2nd prize: Benjamin Moore for improvements to the European Nucleotide Archive article.
3rd prize: Luis Pedro Coelho for improvements to the Bioimage Analysis article.
We are keen to grow the depth and quality of computational biology articles and wish to encourage the widest possible range of students and trainees to take part. We envisage that teachers, tutors, and lecturers could use the competition as an opportunity to train students in literature research on topics of computational biology. This approach to literature review provides the students with a thorough grounding in the subject area of the article. In addition, the collaborative writing environment of Wikipedia encourages critical thinking and improves literature research skills. Furthermore, compared to traditional literature reviews carried out by students, which typically end up unread in a filing cabinet, contributing to Wikipedia means that the students' scholarly contributions will be publicly visible.
We hope that the ISCB Wikipedia competition will continue to grow and help improve the quality of Computational Biology information freely available on the Internet. We are interested in improving not just the articles in Wikipedia, but also the associated media, such as images and figures on Wikimedia Commons, and data through Wikidata. We encourage you to get involved by either entering the competition if you are a student or trainee, or getting your own students to participate. Alex Bateman, Janet Kelso, Daniel Mietchen, Geoff MacIntyre, Tomás Di Domenico, Thomas Abeel, Darren W. Logan, Predrag Radivojac, Burkhard Rost |
PLoS Comput. Biol. | 8 |
| 2012 | Post-translational modifications induce significant yet not extreme changes to protein structureabstractMOTIVATION: A number of studies of individual proteins have shown that post-translational modifications (PTMs) are associated with structural rearrangements of their target proteins. Although such studies provide critical insights into the mechanics behind the dynamic regulation of protein function, they usually feature examples with relatively large conformational changes. However, with the steady growth of Protein Data Bank (PDB) and available PTM sites, it is now possible to more systematically characterize the role of PTMs as conformational switches. In this study, we ask (1) what is the expected extent of structural change upon PTM, (2) how often are those changes in fact substantial, (3) whether the structural impact is spatially localized or global and (4) whether different PTMs have different signatures. RESULTS: We exploit redundancy in PDB and, using root-mean-square deviation, study the conformational heterogeneity of groups of protein structures corresponding to identical sequences in their unmodified and modified forms. We primarily focus on the two most abundant PTMs in PDB, glycosylation and phosphorylation, but show that acetylation and methylation have similar tendencies. Our results provide evidence that PTMs induce conformational changes at both local and global level. However, the proportion of large changes is unexpectedly small; only 7% of glycosylated and 13% of phosphorylated proteins undergo global changes >2 Å. Further analysis suggests that phosphorylation stabilizes protein structure by reducing global conformational heterogeneity by 25%. Overall, these results suggest a subtle but common role of allostery in the mechanisms through which PTMs affect regulatory and signaling pathways. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Fuxiao Xin, Predrag Radivojac |
Bioinform. | 2 |
| 2012 | Computational approaches to protein inference in shotgun proteomicsabstractShotgun proteomics has recently emerged as a powerful approach to characterizing proteomes in biological samples. Its overall objective is to identify the form and quantity of each protein in a high-throughput manner by coupling liquid chromatography with tandem mass spectrometry. As a consequence of its high throughput nature, shotgun proteomics faces challenges with respect to the analysis and interpretation of experimental data. Among such challenges, the identification of proteins present in a sample has been recognized as an important computational task. This task generally consists of (1) assigning experimental tandem mass spectra to peptides derived from a protein database, and (2) mapping assigned peptides to proteins and quantifying the confidence of identified proteins. Protein identification is fundamentally a statistical inference problem with a number of methods proposed to address its challenges. In this review we categorize current approaches into rule-based, combinatorial optimization and probabilistic inference techniques, and present them using integer programming and Bayesian inference frameworks. We also discuss the main challenges of protein identification and propose potential solutions with the goal of spurring innovative research in this area. Yong Fuga Li, Predrag Radivojac |
BMC Bioinform. | 2 |
| 2012 | An Integrated Regulatory Network Reveals Pervasive Cross-Regulation among Transcription and Splicing FactorsabstractTraditionally the gene expression pathway has been regarded as being comprised of independent steps, from RNA transcription to protein translation. To date there is increasing evidence of coupling between the different processes of the pathway, specifically between transcription and splicing. To study the interplay between these processes we derived a transcription-splicing integrated network. The nodes of the network included experimentally verified human proteins belonging to three groups of regulators: transcription factors, splicing factors and kinases. The nodes were wired by instances of predicted transcriptional and alternative splicing regulation. Analysis of the network indicated a pervasive cross-regulation among the nodes; specifically, splicing factors are significantly more connected by alternative splicing regulatory edges relative to the two other subgroups, while transcription factors are more extensively controlled by transcriptional regulation. Furthermore, we found that splicing factors are the most regulated of the three regulatory groups and are subject to extensive combinatorial control by alternative splicing and transcriptional regulation. Consistent with the network results, our bioinformatics analyses showed that the subgroup of kinases have the highest density of predicted phosphorylation sites. Overall, our systematic study reveals that an organizing principle in the logic of integrated networks favor the regulation of regulatory proteins by the specific regulation they conduct. Based on these results, we propose a new regulatory paradigm postulating that gene expression regulation of the master regulators in the cell is predominantly achieved by cross-regulation. Idit Kosti, Predrag Radivojac, Yael Mandel-Gutfreund |
PLoS Comput. Biol. | 2 |
| 2011 | Testing the Ortholog Conjecture with Comparative Functional Genomic Data from MammalsabstractA common assumption in comparative genomics is that orthologous genes share greater functional similarity than do paralogous genes (the "ortholog conjecture"). Many methods used to computationally predict protein function are based on this assumption, even though it is largely untested. Here we present the first large-scale test of the ortholog conjecture using comparative functional genomic data from human and mouse. We use the experimentally derived functions of more than 8,900 genes, as well as an independent microarray dataset, to directly assess our ability to predict function using both orthologs and paralogs. Both datasets show that paralogs are often a much better predictor of function than are orthologs, even at lower sequence identities. Among paralogs, those found within the same species are consistently more functionally similar than those found in a different species. We also find that paralogous pairs residing on the same chromosome are more functionally similar than those on different chromosomes, perhaps due to higher levels of interlocus gene conversion between these pairs. In addition to offering implications for the computational prediction of protein function, our results shed light on the relationship between sequence divergence and functional divergence. We conclude that the most important factor in the evolution of function is not amino acid sequence, but rather the cellular context in which proteins act. Nathan L. Nehrt, Wyatt Travis Clark, Predrag Radivojac, Matthew W. Hahn |
PLoS Comput. Biol. | 3 |
| 2010 | Structure-based kernels for the prediction of catalytic residues and their involvement in human inherited diseaseabstractMOTIVATION: Enzyme catalysis is involved in numerous biological processes and the disruption of enzymatic activity has been implicated in human disease. Despite this, various aspects of catalytic reactions are not completely understood, such as the mechanics of reaction chemistry and the geometry of catalytic residues within active sites. As a result, the computational prediction of catalytic residues has the potential to identify novel catalytic pockets, aid in the design of more efficient enzymes and also predict the molecular basis of disease. RESULTS: We propose a new kernel-based algorithm for the prediction of catalytic residues based on protein sequence, structure and evolutionary information. The method relies upon explicit modeling of similarity between residue-centered neighborhoods in protein structures. We present evidence that this algorithm evaluates favorably against established approaches, and also provides insights into the relative importance of the geometry, physicochemical properties and evolutionary conservation of catalytic residue activity. The new algorithm was used to identify known mutations associated with inherited disease whose molecular mechanism might be predicted to operate specifically though the loss or gain of catalytic residues. It should, therefore, provide a viable approach to identifying the molecular basis of disease in which the loss or gain of function is not caused solely by the disruption of protein stability. Our analysis suggests that both mechanisms are actively involved in human inherited disease. AVAILABILITY AND IMPLEMENTATION: Source code for the structural kernel is available at www.informatics.indiana.edu/predrag/. Fuxiao Xin, Steven Myers, Yong Fuga Li, David N. Cooper, Sean D. Mooney, Predrag Radivojac |
Bioinform. | 6 |
| 2010 | Structure-based kernels for the prediction of catalytic residues and their involvement in human inherited diseaseabstractEnzyme catalysis is involved in numerous biological processes and the disruption of enzymatic activity has been implicated in human disease. Despite the functional importance, various aspects of catalytic reactions are not completely understood, such as the mechanics of reaction chemistry and the geometry of catalytic residues within active sites. As a result, the computational prediction of catalytic residues has the potential to identify novel catalytic pockets, aid in the design of more efficient enzymes and also predict the molecular basis of disease. We proposed a new kernel-based algorithm for the prediction of catalytic residues and functional sites in general in protein structures [ 1 ]. The method relies upon explicit modelling of similarity between residue-centred neighbourhoods in protein structures. Specifically, we start with a construction of oriented structural neighbourhoods followed by separating the neighbourhood volume into small cells. The similarities between two structural neighbourhoods are accumulation of their similarity in each cell. The kernel function is a product of three kernels, each addressing a separate aspect of protein function: (i) the geometric kernel addresses the shape similarity, (ii) the chemical kernel addresses the similarity in physicochemical properties, and (iii) the evolutionary kernel addresses the evolutionary similarity of conservation patterns for the residues in two structural neighbourhoods. Our approach was favourably evaluated against two of the leading alternative approaches, FEATURE [ 2 ] and GBT [ 3 ], as shown in Table 1 . The new algorithm was used to identify known mutations associated with inherited disease whose molecular mechanism might be predicted to operate specifically though the loss or gain of catalytic residues. It should therefore provide a viable approach in identifying the molecular basis of disease in which the loss or gain of function is not caused solely by the disruption of protein stability. Our analysis suggests that both loss and gain of catalytic residues are actively involved in human inherited disease. Our kernel method for functional sites prediction based on protein structures evaluates favourably against established methods on the same data set using the same evaluation procedure. The results from applying our catalytic residue predictor to disease mutations indicated that both loss and gain of catalytic residues are actively involved in human inherited disease. Fuxiao Xin, Steven Myers, Yong Fuga Li, David N. Cooper, Sean D. Mooney, Predrag Radivojac |
BMC Bioinform. | 6 |
| 2009 | Automated inference of molecular mechanisms of disease from amino acid substitutionsabstractMOTIVATION: Advances in high-throughput genotyping and next generation sequencing have generated a vast amount of human genetic variation data. Single nucleotide substitutions within protein coding regions are of particular importance owing to their potential to give rise to amino acid substitutions that affect protein structure and function which may ultimately lead to a disease state. Over the last decade, a number of computational methods have been developed to predict whether such amino acid substitutions result in an altered phenotype. Although these methods are useful in practice, and accurate for their intended purpose, they are not well suited for providing probabilistic estimates of the underlying disease mechanism. RESULTS: We have developed a new computational model, MutPred, that is based upon protein sequence, and which models changes of structural features and functional sites between wild-type and mutant sequences. These changes, expressed as probabilities of gain or loss of structure and function, can provide insight into the specific molecular mechanism responsible for the disease state. MutPred also builds on the established SIFT method but offers improved classification accuracy with respect to human disease mutations. Given conservative thresholds on the predicted disruption of molecular function, we propose that MutPred can generate accurate and reliable hypotheses on the molecular basis of disease for approximately 11% of known inherited disease-causing mutations. We also note that the proportion of changes of functionally relevant residues in the sets of cancer-associated somatic mutations is higher than for the inherited lesions in the Human Gene Mutation Database which are instead predicted to be characterized by disruptions of protein structure. AVAILABILITY: http://mutdb.org/mutpred CONTACT: [email protected]; [email protected]. Vidhya G. Krishnan, Matthew E. Mort, Fuxiao Xin, Kishore K. Kamati, David N. Cooper, Sean D. Mooney, Predrag Radivojac |
Bioinform. | 8 |
| 2009 | Analysis of AML genes in dysregulated molecular networksabstractBACKGROUND: Identifying disease causing genes and understanding their molecular mechanisms are essential to developing effective therapeutics. Thus, several computational methods have been proposed to prioritize candidate disease genes by integrating different data types, including sequence information, biomedical literature, and pathway information. Recently, molecular interaction networks have been incorporated to predict disease genes, but most of those methods do not utilize invaluable disease-specific information available in mRNA expression profiles of patient samples. RESULTS: Through the integration of protein-protein interaction networks and gene expression profiles of acute myeloid leukemia (AML) patients, we identified subnetworks of interacting proteins dysregulated in AML and characterized known mutation genes causally implicated to AML embedded in the subnetworks. The analysis shows that the set of extracted subnetworks is a reservoir rich in AML genes reflecting key leukemogenic processes such as myeloid differentiation, CONCLUSION: We showed that the integrative approach both utilizing gene expression profiles and molecular networks could identify AML causing genes most of which were not detectable with gene expression analysis alone due to their minor changes in mRNA. Hyunchul Jung, Predrag Radivojac, Jong-Won Kim 0001, Doheon Lee |
BMC Bioinform. | 3 |
| 2009 | Influence of Sequence Changes and Environment on Intrinsically Disordered ProteinsabstractMany large-scale studies on intrinsically disordered proteins are implicitly based on the structural models deposited in the Protein Data Bank. Yet, the static nature of deposited models supplies little insight into variation of protein structure and function under diverse cellular and environmental conditions. While the computational predictability of disordered regions provides practical evidence that disorder is an intrinsic property of proteins, the robustness of disordered regions to changes in sequence or environmental conditions has not been systematically studied. We analyzed intrinsically disordered regions in the same or similar proteins crystallized independently and studied their sensitivity to changes in protein sequence and parameters of crystallographic experiments. The observed changes in the existence, position, and length of disordered regions indicate that their appearance in X-ray structures dramatically depends on changes in amino acid sequence and peculiarities of the crystallographic experiment. Our study also raises general questions regarding protein evolution and the regulation of protein structure, dynamics, and function via variations in cellular and environmental conditions. Amrita Mohan, Vladimir N. Uversky, Predrag Radivojac |
PLoS Comput. Biol. | 3 |
| 2008 | A Bayesian Approach to Protein Inference Problem in Shotgun Proteomics
Yong Fuga Li, Randy J. Arnold, Predrag Radivojac, Quanhu Sheng, Haixu Tang |
RECOMB | 4 |
| 2008 | Fast and accurate identification of semi-tryptic peptides in shotgun proteomicsabstractMOTIVATION: One of the major problems in shotgun proteomics is the low peptide coverage when analyzing complex protein samples. Identifying more peptides, e.g. non-tryptic peptides, may increase the peptide coverage and improve protein identification and/or quantification that are based on the peptide identification results. Searching for all potential non-tryptic peptides is, however, time consuming for shotgun proteomics data from complex samples, and poses a challenge for a routine data analysis. RESULTS: We hypothesize that non-tryptic peptides are mainly created from the truncation of regular tryptic peptides before separation. We introduce the notion of truncatability of a tryptic peptide, i.e. the probability of the peptide to be identified in its truncated form, and build a predictor to estimate a peptide's truncatability from its sequence. We show that our predictions achieve useful accuracy, with the area under the ROC curve from 76% to 87%, and can be used to filter the sequence database for identifying truncated peptides. After filtering, only a limited number of tryptic peptides with the highest truncatability are retained for non-tryptic peptide searching. By applying this method to identification of semi-tryptic peptides, we show that a significant number of such peptides can be identified within a searching time comparable to that of tryptic peptide identification. Pedro Alves, Randy J. Arnold, David E. Clemmer, James P. Reilly, Quanhu Sheng, Haixu Tang, Zhiyin Xun, Predrag Radivojac |
Bioinform. | 10 |
| 2006 | Using Compression to Identify Classes of Inauthentic TextsabstractRecent events have made it clear that some kinds of technical texts, generated by machine and essentially meaningless, can be confused with authentic, technical texts written by humans. We identify this as a potential problem, since no existing systems for, say the web, can or do discriminate on this basis. We believe that there are subtle, short- and long-range word or even string repetitions extant in human texts, but not in many classes of computer generated texts, that can be used to discriminate based on meaning. In this paper we employ universal lossless source coding to generate features in a high-dimensional space and then apply support vector machines to discriminate between the classes of authentic and inauthentic expository texts. Compression profiles for the two kinds of text are distinct—the authentic texts being bounded by various classes of more compressible or less compressible texts that are computer generated. This in turn led to the high prediction accuracy of our models which support a conjecture that there exists a relationship between meaning and compressibility. Our results show that the learning algorithm based upon the compression profile outperformed standard term-frequency text categorization on several non-trivial classes of inauthentic texts. Availability: http://www.informatics.indiana.edu/predrag/fsi.htm. Mehmet M. Dalkilic, Wyatt Travis Clark, James C. Costello, Predrag Radivojac |
SDM | 4 |
| 2006 | Two Sample Logo: a graphical representation of the differences between two sets of sequence alignmentsabstractSUMMARY: Two Sample Logo is a web-based tool that detects and displays statistically significant differences in position-specific symbol compositions between two sets of multiple sequence alignments. In a typical scenario, two groups of aligned sequences will share a common motif but will differ in their functional annotation. The inclusion of the background alignment provides an appropriate underlying amino acid or nucleotide distribution and addresses intersite symbol correlations. In addition, the difference detection process is sensitive to the sizes of the aligned groups. Two Sample Logo extends WebLogo, a widely-used sequence logo generator. The source code is distributed under the MIT Open Source license agreement and is available for download free of charge. Vladimir Vacic, Lilia M. Iakoucheva, Predrag Radivojac |
Bioinform. | 3 |
| 2006 | Length-dependent prediction of protein intrinsic disorderabstractBACKGROUND: Due to the functional importance of intrinsically disordered proteins or protein regions, prediction of intrinsic protein disorder from amino acid sequence has become an area of active research as witnessed in the 6th experiment on Critical Assessment of Techniques for Protein Structure Prediction (CASP6). Since the initial work by Romero et al. (Identifying disordered regions in proteins from amino acid sequences, IEEE Int. Conf. Neural Netw., 1997), our group has developed several predictors optimized for long disordered regions (>30 residues) with prediction accuracy exceeding 85%. However, these predictors are less successful on short disordered regions (< or =30 residues). A probable cause is a length-dependent amino acid compositions and sequence properties of disordered regions. RESULTS: We proposed two new predictor models, VSL2-M1 and VSL2-M2, to address this length-dependency problem in prediction of intrinsic protein disorder. These two predictors are similar to the original VSL1 predictor used in the CASP6 experiment. In both models, two specialized predictors were first built and optimized for short (< or = 30 residues) and long disordered regions (>30 residues), respectively. A meta predictor was then trained to integrate the specialized predictors into the final predictor model. As the 10-fold cross-validation results showed, the VSL2 predictors achieved well-balanced prediction accuracies of 81% on both short and long disordered regions. Comparisons over the VSL2 training dataset via 10-fold cross-validation and a blind-test set of unrelated recent PDB chains indicated that VSL2 predictors were significantly more accurate than several existing predictors of intrinsic protein disorder. CONCLUSION: The VSL2 predictors are applicable to disordered regions of any length and can accurately identify the short disordered regions that are often misclassified by our previous disorder predictors. The success of the VSL2 predictors further confirmed the previously observed differences in amino acid compositions and sequence properties between short and long disordered regions, and justified our approaches for modelling short and long disordered regions separately. The VSL2 predictors are freely accessible for non-commercial use at http://www.ist.temple.edu/disprot/predictorVSL2.php. Kang Peng, Predrag Radivojac, Slobodan Vucetic, A. Keith Dunker, Zoran Obradovic |
BMC Bioinform. | 2 |
| 2006 | Intrinsic Disorder Is a Common Feature of Hub Proteins from Four Eukaryotic InteractomesabstractRecent proteome-wide screening approaches have provided a wealth of information about interacting proteins in various organisms. To test for a potential association between protein connectivity and the amount of predicted structural disorder, the disorder propensities of proteins with various numbers of interacting partners from four eukaryotic organisms (Caenorhabditis elegans, Saccharomyces cerevisiae, Drosophila melanogaster, and Homo sapiens) were investigated. The results of PONDR VL-XT disorder analysis show that for all four studied organisms, hub proteins, defined here as those that interact with > or = 10 partners, are significantly more disordered than end proteins, defined here as those that interact with just one partner. The proportion of predicted disordered residues, the average disorder score, and the number of predicted disordered regions of various lengths were higher overall in hubs than in ends. A binary classification of hubs and ends into ordered and disordered subclasses using the consensus prediction method showed a significant enrichment of wholly disordered proteins and a significant depletion of wholly ordered proteins in hubs relative to ends in worm, fly, and human. The functional annotation of yeast hubs and ends using GO categories and the correlation of these annotations with disorder predictions demonstrate that proteins with regulation, transcription, and development annotations are enriched in disorder, whereas proteins with catalytic activity, transport, and membrane localization annotations are depleted in disorder. The results of this study demonstrate that intrinsic structural disorder is a distinctive and common characteristic of eukaryotic hub proteins, and that disorder may serve as a determinant of protein interactivity. Chad Haynes, Christopher J. Oldfield, Niels Klitgord, Michael E. Cusick, Predrag Radivojac, Vladimir N. Uversky, Marc Vidal, Lilia M. Iakoucheva |
PLoS Comput. Biol. | 6 |
| 2005 | Intrinsic Disorder and Protein modifications: Building an SVM Predictor for Methylation
Kenneth Daily, Predrag Radivojac, A. Keith Dunker |
CIBCB | 2 |
| 2005 | DisProt: a database of protein disorderabstractUNLABELLED: The Database of Protein Disorder (DisProt) is a curated database that provides structure and function information about proteins that lack a fixed three-dimensional (3D) structure under putatively native conditions, either in their entirety or in part. Starting from the central premise that intrinsic disorder is an important structural class of protein and in order to meet the increasing interest thereof, DisProt is aimed at becoming a central repository of disorder-related information. For each disordered protein, the database includes the name of the protein, various aliases, accession codes, amino acid sequence, location of the disordered region(s), and methods used for structural (disorder) characterization. If applicable, most entries also list the biological function(s) of each disordered region, how each region of disorder is used for function, as well as provide links to PubMed abstracts and major protein databases. AVAILABILITY: www.disprot.org Slobodan Vucetic, Zoran Obradovic, Vladimir Vacic, Predrag Radivojac, Kang Peng, Lilia M. Iakoucheva, Marc S. Cortese, J. David Lawson, Celeste J. Brown, Jason G. Sikes, Crystal D. Newton, A. Keith Dunker |
Bioinform. | 4 |
| 2004 | Feature Selection Filters Based on the Permutation Test
Predrag Radivojac, Zoran Obradovic, A. Keith Dunker, Slobodan Vucetic |
ECML | 1 |
| 2004 | Classification and knowledge discovery in protein databases
Predrag Radivojac, Nitesh V. Chawla, A. Keith Dunker, Zoran Obradovic |
J. Biomed. Informatics | 1 |