VLDB 2026 Research / reviewers in the wild / expert
Sanjay Joshua Swamidass
dblp:79/3967 · also S. Joshua Swamidass
· DBLP profile ↗
22ranked-venue papers
3as first author
4since 2021 · last 2022
0000-0003-2191-0778ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 18 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
11 papers |
Bioinformatics and computational biology · 91% Computational science and engineering · 9% |
Topics — the 19 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › drug discovery
drug metabolism prediction |
0.9 | 4 | 2016 | A simple model predicts UGT-mediated metabolism · Bioinform. 2016 Extending P450 site-of-metabolism models with region-resolution data · Bioinform. 2015 XenoSite server: a web-available site of metabolism prediction tool · Bioinform. 2015 |
Bioinformatics and computational biology › drug discovery › drug metabolism prediction
site of metabolism prediction |
0.6 | 3 | 2015 | Extending P450 site-of-metabolism models with region-resolution data · Bioinform. 2015 XenoSite server: a web-available site of metabolism prediction tool · Bioinform. 2015 RS-WebPredictor: a server for predicting CYP-mediated sites of metabolism on drug-like molecules · Bioinform. 2013 |
Computational science and engineering
computational chemistry |
0.4 | 2 | 2015 | Extending P450 site-of-metabolism models with region-resolution data · Bioinform. 2015 RS-WebPredictor: a server for predicting CYP-mediated sites of metabolism on drug-like molecules · Bioinform. 2013 |
Bioinformatics and computational biology › sequence analysis › sequence motif analysis
position weight matrix |
0.3 | 1 | 2017 | BEESEM: estimation of binding energy models using HT-SELEX data · Bioinform. 2017 |
Bioinformatics and computational biology › sequence analysis › motif discovery
transcription factor binding motif discovery |
0.3 | 1 | 2017 | BEESEM: estimation of binding energy models using HT-SELEX data · Bioinform. 2017 |
Bioinformatics and computational biology › gene regulation › transcription factor binding
transcription factor binding specificity |
0.3 | 1 | 2017 | BEESEM: estimation of binding energy models using HT-SELEX data · Bioinform. 2017 |
Bioinformatics and computational biology › drug discovery
computational drug discovery |
0.2 | 2 | 2011 | Enhancing the rate of scaffold discovery with diversity-oriented prioritization · Bioinform. 2011 A CROC stronger than ROC: measuring, visualizing and optimizing early retrieval · Bioinform. 2010 |
Bioinformatics and computational biology › molecular informatics
cheminformatics |
0.2 | 2 | 2013 | Scaffold network generator: a tool for mining molecular structures · Bioinform. 2013 ChemDB: a public database of small molecules and related chemoinformatics resources · Bioinform. 2005 |
Bioinformatics and computational biology › cancer genomics
cancer driver gene identification |
0.2 | 1 | 2015 | Statistically identifying tumor suppressors and oncogenes from pan-cancer genome-sequencing data · Bioinform. 2015 |
Bioinformatics and computational biology › sequence analysis › sequence assembly
scaffold graph analysis |
0.2 | 1 | 2013 | Scaffold network generator: a tool for mining molecular structures · Bioinform. 2013 |
Bioinformatics and computational biology › drug discovery
high-throughput screening |
0.1 | 1 | 2011 | Enhancing the rate of scaffold discovery with diversity-oriented prioritization · Bioinform. 2011 |
Bioinformatics and computational biology › drug discovery › high-throughput screening
hit selection |
0.1 | 1 | 2011 | Enhancing the rate of scaffold discovery with diversity-oriented prioritization · Bioinform. 2011 |
Bioinformatics and computational biology
machine learning for biology |
0.1 | 1 | 2010 | A CROC stronger than ROC: measuring, visualizing and optimizing early retrieval · Bioinform. 2010 |
Bioinformatics and computational biology › drug discovery
virtual screening |
0.1 | 1 | 2010 | A CROC stronger than ROC: measuring, visualizing and optimizing early retrieval · Bioinform. 2010 |
Bioinformatics and computational biology › drug discovery › molecular optimization
lead optimization |
0.1 | 1 | 2016 | A simple model predicts UGT-mediated metabolism · Bioinform. 2016 |
Bioinformatics and computational biology › molecular informatics › cheminformatics
chemical database |
0.1 | 1 | 2007 | ChemDB update - full-text search and virtual chemical space · Bioinform. 2007 |
Bioinformatics and computational biology › molecular informatics › cheminformatics
chemical structure search |
0.1 | 1 | 2007 | ChemDB update - full-text search and virtual chemical space · Bioinform. 2007 |
Bioinformatics and computational biology › cancer genomics
somatic mutation analysis |
0.1 | 1 | 2015 | Statistically identifying tumor suppressors and oncogenes from pan-cancer genome-sequencing data · Bioinform. 2015 |
Bioinformatics and computational biology
molecular property prediction |
0.0 | 1 | 2005 | ChemDB: a public database of small molecules and related chemoinformatics resources · Bioinform. 2005 |
Methods — techniques the papers use, named apart from their topics
expectation-maximization · 0.5machine learning · 0.4biophysical modeling · 0.3HT-SELEX · 0.3neural network · 0.2heuristic model · 0.2statistical hypothesis testing · 0.2random forest · 0.2patient bias signal · 0.2scaffold network generation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Fair-Net: A Network Architecture for Reducing Performance Disparity between Identifiable Sub-populationsabstractIn real world datasets, particular groups are under-represented, much rarer than others, and machine learning classifiers will often preform worse on under-represented populations. This problem is aggravated across many domains where datasets are class imbalanced, with a minority class far rarer than the majority class. Naive approaches to handle under-representation and class imbalance include training sub-population specific classifiers that handle class imbalance or training a global classifier that overlooks sub-population disparities and aims to achieve high overall accuracy by handling class imbalance. In this study, we find that these approaches are vulnerable in class imbalanced datasets with minority sub-populations. We introduced Fair-Net, a branched multitask neural network architecture that improves both classification accuracy and probability calibration across identifiable sub-populations in class imbalanced datasets. Fair-Nets is a straightforward extension to the output layer and error function of a network, so can be incorporated in far more complex architectures. Empirical studies with three real world benchmark datasets demonstrate that Fair-Net improves classification and calibration performance, substantially reducing performance disparity between gender and racial sub-populations. Arghya Datta, Sanjay Joshua Swamidass |
ICAART (3) | 2 |
| 2021 | Cal-Net: Jointly Learning Classification and Calibration On Imbalanced Binary Classification TasksabstractDatasets in critical domains are often class imbalanced, with a minority class far rarer than the majority class, and classification models face challenges to produce calibrated predictions on these datasets. A common approach to address this issue is to train classification models in the first step and subsequently use post-processing parametric or non-parametric calibration techniques to re-scale the model's outputs in the second step without tuning any underlying parameters in the model to improve calibration. In this study, we have shown that these common approaches are vulnerable to class imbalanced data, often producing unstable results that do not jointly optimize classification or calibration performance. We have introduced Cal-Net, a “self-calibrating” neural network architecture that simultaneously optimizes classification and calibration performances for class imbalanced datasets in a single training phase, thereby eliminating the need for any post-processing procedure for confidence calibration. Empirical results have shown that Cal-Net outperforms far more complex neural networks and post-processing calibration techniques in both classification and calibration performances on four synthetic and four benchmark class imbalanced binary classification datasets. Furthermore, Cal-Net can readily be extended to more complicated learning tasks, online learning and can be incorporated in more complex architectures as the final state. Arghya Datta, Noah R. Flynn, Sanjay Joshua Swamidass |
IJCNN | 3 |
| 2021 | Machine learning liver-injuring drug interactions with non-steroidal anti-inflammatory drugs (NSAIDs) from a retrospective electronic health record (EHR) cohortabstractDrug-drug interactions account for up to 30% of adverse drug reactions. Increasing prevalence of electronic health records (EHRs) offers a unique opportunity to build machine learning algorithms to identify drug-drug interactions that drive adverse events. In this study, we investigated hospitalizations' data to study drug interactions with non-steroidal anti-inflammatory drugs (NSAIDS) that result in drug-induced liver injury (DILI). We propose a logistic regression based machine learning algorithm that unearths several known interactions from an EHR dataset of about 400,000 hospitalization. Our proposed modeling framework is successful in detecting 87.5% of the positive controls, which are defined by drugs known to interact with diclofenac causing an increased risk of DILI, and correctly ranks aggregate risk of DILI for eight commonly prescribed NSAIDs. We found that our modeling framework is particularly successful in inferring associations of drug-drug interactions from relatively small EHR datasets. Furthermore, we have identified a novel and potentially hepatotoxic interaction that might occur during concomitant use of meloxicam and esomeprazole, which are commonly prescribed together to allay NSAID-induced gastrointestinal (GI) bleeding. Empirically, we validate our approach against prior methods for signal detection on EHR datasets, in which our proposed approach outperforms all the compared methods across most metrics, such as area under the receiver operating characteristic curve (AUROC) and area under the precision-recall curve (AUPRC). Arghya Datta, Noah R. Flynn, Dustyn A. Barnette, Keith F. Woeltje, Grover P. Miller, Sanjay Joshua Swamidass |
PLoS Comput. Biol. | 6 |
| 2021 | 'Black Box' to 'Conversational' Machine Learning: Ondansetron Reduces Risk of Hospital-Acquired Venous ThromboembolismabstractMachine learning, combined with a proliferation of electronic healthcare records (EHR), has the potential to transform medicine by identifying previously unknown interventions that reduce the risk of adverse outcomes. To realize this potential, machine learning must leave the conceptual 'black box' in complex domains to overcome several pitfalls, like the presence of confounding variables. These variables predict outcomes but are not causal, often yielding uninformative models. In this work, we envision a 'conversational' approach to design machine learning models, which couple modeling decisions to domain expertise. We demonstrate this approach via a retrospective cohort study to identify factors which affect the risk of hospital-acquired venous thromboembolism (HA-VTE). Using logistic regression for modeling, we have identified drugs that reduce the risk of HA-VTE. Our analysis reveals that ondansetron, an anti-nausea and anti-emetic medication, commonly used in treating side-effects of chemotherapy and post-general anesthesia period, substantially reduces the risk of HA-VTE when compared to aspirin (11% vs. 15% relative risk reduction or RRR, respectively). The low cost and low morbidity of ondansetron may justify further inquiry into its use as a preventative agent for HA-VTE. This case study highlights the importance of engaging domain expertise while applying machine learning in complex domains. Arghya Datta, Matthew K. Matlock, Na Le Dang, Thiago C. Moulin, Keith F. Woeltje, Elizabeth L. Yanik, Sanjay Joshua Swamidass |
IEEE J. Biomed. Health Informatics | 7 |
| 2019 | Deep learning long-range information in undirected graphs with wave networksabstractGraph algorithms are key tools in many fields of science and technology. Some of these algorithms depend on propagating information between distant nodes in a graph. Recently, there have been a number of deep learning architectures proposed to learn on undirected graphs. However, most of these architectures aggregate information in the local neighborhood of a node, and therefore they may not be capable of efficiently propagating long-range information. To solve this problem we examine a recently proposed architecture, wave, which propagates information back and forth across an undirected graph in waves of nonlinear computation. We compare wave to graph convolution, an architecture based on local aggregation, and find that wave learns three different graph-based tasks with greater efficiency and accuracy. These three tasks include (1) labeling a path connecting two nodes in a graph, (2) solving a maze presented as an image, and (3) computing voltages in a circuit. These tasks range from trivial to very difficult, but wave can extrapolate from small training examples to much larger testing examples. These results show that wave may be able to efficiently solve a wide range of tasks that require long-range information propagation across undirected graphs. An implementation of the wave network, and example code for the maze task are included in the tflon deep learning toolkit (https://bitbucket.org/mkmatlock/tflon). Matthew K. Matlock, Arghya Datta, Na Le Dang, Kevin Jiang, Sanjay Joshua Swamidass |
IJCNN | 5 |
| 2018 | Deep Learning Global Glomerulosclerosis in Transplant Kidney Frozen SectionsabstractTransplantable kidneys are in very limited supply. Accurate viability assessment prior to transplantation could minimize organ discard. Rapid and accurate evaluation of intra-operative donor kidney biopsies is essential for determining which kidneys are eligible for transplantation. The criterion for accepting or rejecting donor kidneys relies heavily on pathologist determination of the percent of glomeruli (determined from a frozen section) that are normal and sclerotic. This percentage is a critical measurement that correlates with transplant outcome. Inter- and intra-observer variability in donor biopsy evaluation is, however, significant. An automated method for determination of percent global glomerulosclerosis could prove useful in decreasing evaluation variability, increasing throughput, and easing the burden on pathologists. Here, we describe the development of a deep learning model that identifies and classifies non-sclerosed and sclerosed glomeruli in whole-slide images of donor kidney frozen section biopsies. This model extends a convolutional neural network (CNN) pre-trained on a large database of digital images. The extended model, when trained on just 48 whole slide images, exhibits slide-level evaluation performance on par with expert renal pathologists. Encouragingly, the model's performance is robust to slide preparation artifacts associated with frozen section preparation. The model substantially outperforms a model trained on image patches of isolated glomeruli, in terms of both accuracy and speed. The methodology overcomes the technical challenge of applying a pretrained CNN bottleneck model to whole-slide image classification. The traditional patch-based approach, while exhibiting deceptively good performance classifying isolated patches, does not translate successfully to whole-slide image segmentation in this setting. As the first model reported that identifies and classifies normal and sclerotic glomeruli in frozen kidney sections, and thus the first model reported in the literature relevant to kidney transplantation, it may become an essential part of donor kidney biopsy evaluation in the clinical setting. Jon N. Marsh, Matthew K. Matlock, Satoru Kudose, Ta-Chiang Liu, Thaddeus S. Stappenbeck, Joseph P. Gaut, Sanjay Joshua Swamidass |
IEEE Trans. Medical Imaging | 7 |
| 2017 | BEESEM: estimation of binding energy models using HT-SELEX dataabstractMOTIVATION: Characterizing the binding specificities of transcription factors (TFs) is crucial to the study of gene expression regulation. Recently developed high-throughput experimental methods, including protein binding microarrays (PBM) and high-throughput SELEX (HT-SELEX), have enabled rapid measurements of the specificities for hundreds of TFs. However, few studies have developed efficient algorithms for estimating binding motifs based on HT-SELEX data. Also the simple method of constructing a position weight matrix (PWM) by comparing the frequency of the preferred sequence with single-nucleotide variants has the risk of generating motifs with higher information content than the true binding specificity. RESULTS: We developed an algorithm called BEESEM that builds on a comprehensive biophysical model of protein-DNA interactions, which is trained using the expectation maximization method. BEESEM is capable of selecting the optimal motif length and calculating the confidence intervals of estimated parameters. By comparing BEESEM with the published motifs estimated using the same HT-SELEX data, we demonstrate that BEESEM provides significant improvements. We also evaluate several motif discovery algorithms on independent PBM and ChIP-seq data. BEESEM provides significantly better fits to in vitro data, but its performance is similar to some other methods on in vivo data under the criterion of the area under the receiver operating characteristic curve (AUROC). This highlights the limitations of the purely rank-based AUROC criterion. Using quantitative binding data to assess models, however, demonstrates that BEESEM improves on prior models. AVAILABILITY AND IMPLEMENTATION: Freely available on the web at http://stormo.wustl.edu/resources.html . CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shuxiang Ruan, Sanjay Joshua Swamidass, Gary D. Stormo |
Bioinform. | 2 |
| 2016 | A survey of current trends in computational drug repositioningabstractComputational drug repositioning or repurposing is a promising and efficient tool for discovering new uses from existing drugs and holds the great potential for precision medicine in the age of big data. The explosive growth of large-scale genomic and phenotypic data, as well as data of small molecular compounds with granted regulatory approval, is enabling new developments for computational repositioning. To achieve the shortest path toward new drug indications, advanced data processing and analysis strategies are critical for making sense of these heterogeneous molecular measurements. In this review, we show recent advancements in the critical areas of computational drug repositioning from multiple aspects. First, we summarize available data sources and the corresponding computational repositioning strategies. Second, we characterize the commonly used computational techniques. Third, we discuss validation strategies for repositioning studies, including both computational and experimental methods. Finally, we highlight potential opportunities and use-cases, including a few target areas such as cancers. We conclude with a brief discussion of the remaining challenges in computational drug repositioning. Jiao Li 0001, Si Zheng 0001, Atul J. Butte, Sanjay Joshua Swamidass, Zhiyong Lu |
Briefings Bioinform. | 5 |
| 2016 | A simple model predicts UGT-mediated metabolismabstractMOTIVATION: Uridine diphosphate glucunosyltransferases (UGTs) metabolize 15% of FDA approved drugs. Lead optimization efforts benefit from knowing how candidate drugs are metabolized by UGTs. This paper describes a computational method for predicting sites of UGT-mediated metabolism on drug-like molecules. RESULTS: XenoSite correctly predicts test molecule's sites of glucoronidation in the Top-1 or Top-2 predictions at a rate of 86 and 97%, respectively. In addition to predicting common sites of UGT conjugation, like hydroxyl groups, it can also accurately predict the glucoronidation of atypical sites, such as carbons. We also describe a simple heuristic model for predicting UGT-mediated sites of metabolism that performs nearly as well (with, respectively, 80 and 91% Top-1 and Top-2 accuracy), and can identify the most challenging molecules to predict on which to assess more complex models. Compared with prior studies, this model is more generally applicable, more accurate and simpler (not requiring expensive quantum modeling). AVAILABILITY AND IMPLEMENTATION: The UGT metabolism predictor developed in this study is available at http://swami.wustl.edu/xenosite/p/ugt CONTACT: : [email protected] information: Supplementary data are available at Bioinformatics online. Na Le Dang, Tyler B. Hughes, Varun Krishnamurthy, Sanjay Joshua Swamidass |
Bioinform. | 4 |
| 2015 | Statistically identifying tumor suppressors and oncogenes from pan-cancer genome-sequencing dataabstractMOTIVATION: Several tools exist to identify cancer driver genes based on somatic mutation data. However, these tools do not account for subclasses of cancer genes: oncogenes, which undergo gain-of-function events, and tumor suppressor genes (TSGs) which undergo loss-of-function. A method which accounts for these subclasses could improve performance while also suggesting a mechanism of action for new putative cancer genes. RESULTS: We develop a panel of five complementary statistical tests and assess their performance against a curated set of 99 HiConf cancer genes using a pan-cancer dataset of 1.7 million mutations. We identify patient bias as a novel signal for cancer gene discovery, and use it to significantly improve detection of oncogenes over existing methods (AUROC = 0.894). Additionally, our test of truncation event rate separates oncogenes and TSGs from one another (AUROC = 0.922). Finally, a random forest integrating the five tests further improves performance and identifies new cancer genes, including CACNG3, HDAC2, HIST1H1E, NXF1, GPS2 and HLA-DRB1. AVAILABILITY AND IMPLEMENTATION: All mutation data, instructions, functions for computing the statistics and integrating them, as well as the HiConf gene panel, are available at www.github.com/Bose-Lab/Improved-Detection-of-Cancer-Genes. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Runjun D. Kumar, Adam C. Searleman, Sanjay Joshua Swamidass, Obi L. Griffith, Ron Bose |
Bioinform. | 3 |
| 2015 | XenoSite server: a web-available site of metabolism prediction toolabstractUNLABELLED: Cytochrome P450 enzymes (P450s) are metabolic enzymes that process the majority of FDA-approved, small-molecule drugs. Understanding how these enzymes modify molecule structure is key to the development of safe, effective drugs. XenoSite server is an online implementation of the XenoSite, a recently published computational model for P450 metabolism. XenoSite predicts which atomic sites of a molecule--sites of metabolism (SOMs)--are modified by P450s. XenoSite server accepts input in common chemical file formats including SDF and SMILES and provides tools for visualizing the likelihood that each atomic site is a site of metabolism for a variety of important P450s, as well as a flat file download of SOM predictions. AVAILABILITY AND IMPLEMENTATION: XenoSite server is available at http://swami.wustl.edu/xenosite. Matthew K. Matlock, Tyler B. Hughes, Sanjay Joshua Swamidass |
Bioinform. | 3 |
| 2015 | Extending P450 site-of-metabolism models with region-resolution dataabstractMOTIVATION: Cytochrome P450s are a family of enzymes responsible for the metabolism of approximately 90% of FDA-approved drugs. Medicinal chemists often want to know which atoms of a molecule-its metabolized sites-are oxidized by Cytochrome P450s in order to modify their metabolism. Consequently, there are several methods that use literature-derived, atom-resolution data to train models that can predict a molecule's sites of metabolism. There is, however, much more data available at a lower resolution, where the exact site of metabolism is not known, but the region of the molecule that is oxidized is known. Until now, no site-of-metabolism models made use of region-resolution data. RESULTS: Here, we describe XenoSite-Region, the first reported method for training site-of-metabolism models with region-resolution data. Our approach uses the Expectation Maximization algorithm to train a site-of-metabolism model. Region-resolution metabolism data was simulated from a large site-of-metabolism dataset, containing 2000 molecules with 3400 metabolized and 30 000 un-metabolized sites and covering nine Cytochrome P450 isozymes. When training on the same molecules (but with only region-level information), we find that this approach yields models almost as accurate as models trained with atom-resolution data. Moreover, we find that atom-resolution trained models are more accurate when also trained with region-resolution data from additional molecules. Our approach, therefore, opens up a way to extend the applicable domain of site-of-metabolism models into larger regions of chemical space. This meets a critical need in drug development by tapping into underutilized data commonly available in most large drug companies. AVAILABILITY AND IMPLEMENTATION: The algorithm, data and a web server are available at http://swami.wustl.edu/xregion. Jed Zaretzki, Michael R. Browning, Tyler B. Hughes, Sanjay Joshua Swamidass |
Bioinform. | 4 |
| 2013 | Accounting for noise when clustering biological dataabstractClustering is a powerful and commonly used technique that organizes and elucidates the structure of biological data. Clustering data from gene expression, metabolomics and proteomics experiments has proven to be useful at deriving a variety of insights, such as the shared regulation or function of biochemical components within networks. However, experimental measurements of biological processes are subject to substantial noise-stemming from both technical and biological variability-and most clustering algorithms are sensitive to this noise. In this article, we explore several methods of accounting for noise when analyzing biological data sets through clustering. Using a toy data set and two different case studies-gene expression and protein phosphorylation-we demonstrate the sensitivity of clustering algorithms to noise. Several methods of accounting for this noise can be used to establish when clustering results can be trusted. These methods span a range of assumptions about the statistical properties of the noise and can therefore be applied to virtually any biological data source. Roman Sloutsky, Nicolas Jimenez, Sanjay Joshua Swamidass, Kristen M. Naegle |
Briefings Bioinform. | 3 |
| 2013 | Scaffold network generator: a tool for mining molecular structuresabstractSUMMARY: Scaffold network generator (SNG) is an open-source command-line utility that computes the hierarchical network of scaffolds that define a large set of input molecules. Scaffold networks are useful for visualizing, analysing and understanding the chemical data that is increasingly available through large public repositories like PubChem. For example, some groups have used scaffold networks to identify missed-actives in high-throughput screens of small molecules with bioassays. Substantially improving on existing software, SNG is robust enough to work on millions of molecules at a time with a simple command-line interface. AVAILABILITY AND IMPLEMENTATION: SNG is accessible at http://swami.wustl.edu/sng Matthew K. Matlock, Jed Zaretzki, Sanjay Joshua Swamidass |
Bioinform. | 3 |
| 2013 | RS-WebPredictor: a server for predicting CYP-mediated sites of metabolism on drug-like moleculesabstractSUMMARY: Regioselectivity-WebPredictor (RS-WebPredictor) is a server that predicts isozyme-specific cytochrome P450 (CYP)-mediated sites of metabolism (SOMs) on drug-like molecules. Predictions may be made for the promiscuous 2C9, 2D6 and 3A4 CYP isozymes, as well as CYPs 1A2, 2A6, 2B6, 2C8, 2C19 and 2E1. RS-WebPredictor is the first freely accessible server that predicts the regioselectivity of the last six isozymes. Server execution time is fast, taking on average 2s to encode a submitted molecule and 1s to apply a given model, allowing for high-throughput use in lead optimization projects. AVAILABILITY: RS-WebPredictor is accessible for free use at http://reccr.chem.rpi.edu/Software/RS-WebPredictor/ Jed Zaretzki, Charles Bergeron, Tao-wei Huang, Patrik Rydberg, Sanjay Joshua Swamidass, Curt M. Breneman |
Bioinform. | 5 |
| 2011 | Mining small-molecule screens to repurpose drugsabstractRepurposing and repositioning drugs--discovering new uses for existing and experimental medicines-is an attractive strategy for rescuing stalled pharmaceutical projects, finding treatments for neglected diseases, and reducing the time, cost and risk of drug development. As this strategy emerged, academic researchers began performing high-throughput screens (HTS) of small molecules--the type of experiments once exclusively conducted in industry--and making the data from these screens available to all. Several methods can mine this data to inform repurposing and repositioning efforts. Despite these methods' limitations, it is hopeful that they will accelerate the discovery of new uses for known drugs, but this hope has not yet been realized. Sanjay Joshua Swamidass |
Briefings Bioinform. | 1 |
| 2011 | Enhancing the rate of scaffold discovery with diversity-oriented prioritizationabstractMOTIVATION: In high-throughput screens (HTS) of small molecules for activity in an in vitro assay, it is common to search for active scaffolds, with at least one example successfully confirmed as an active. The number of active scaffolds better reflects the success of the screen than the number of active molecules. Many existing algorithms for deciding which hits should be sent for confirmatory testing neglect this concern. RESULTS: We derived a new extension of a recently proposed economic framework, diversity-oriented prioritization (DOP), that aims-by changing which hits are sent for confirmatory testing-to maximize the number of scaffolds with at least one confirmed active. In both retrospective and prospective experiments, DOP accurately predicted the number of scaffold discoveries in a batch of confirmatory experiments, improved the rate of scaffold discovery by 8-17%, and was surprisingly robust to the size of the confirmatory test batches. As an extension of our previously reported economic framework, DOP can be used to decide the optimal number of hits to send for confirmatory testing by iteratively computing the cost of discovering an additional scaffold, the marginal cost of discovery. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Sanjay Joshua Swamidass, Bradley T. Calhoun, Joshua A. Bittker, Nicole E. Bodycombe, Paul A. Clemons |
Bioinform. | 1 |
| 2010 | A CROC stronger than ROC: measuring, visualizing and optimizing early retrievalabstractMOTIVATION: The performance of classifiers is often assessed using Receiver Operating Characteristic ROC [or (AC) accumulation curve or enrichment curve] curves and the corresponding areas under the curves (AUCs). However, in many fundamental problems ranging from information retrieval to drug discovery, only the very top of the ranked list of predictions is of any interest and ROCs and AUCs are not very useful. New metrics, visualizations and optimization tools are needed to address this 'early retrieval' problem. RESULTS: To address the early retrieval problem, we develop the general concentrated ROC (CROC) framework. In this framework, any relevant portion of the ROC (or AC) curve is magnified smoothly by an appropriate continuous transformation of the coordinates with a corresponding magnification factor. Appropriate families of magnification functions confined to the unit square are derived and their properties are analyzed together with the resulting CROC curves. The area under the CROC curve (AUC[CROC]) can be used to assess early retrieval. The general framework is demonstrated on a drug discovery problem and used to discriminate more accurately the early retrieval performance of five different predictors. From this framework, we propose a novel metric and visualization-the CROC(exp), an exponential transform of the ROC curve-as an alternative to other methods. The CROC(exp) provides a principled, flexible and effective way for measuring and visualizing early retrieval performance with excellent statistical power. Corresponding methods for optimizing early retrieval are also described in the Appendix. AVAILABILITY: Datasets are publicly available. Python code and command-line utilities implementing CROC curves and metrics are available at http://pypi.python.org/pypi/CROC/ CONTACT: [email protected] Sanjay Joshua Swamidass, Chloé-Agathe Azencott, Kenneth Daily, Pierre Baldi |
Bioinform. | 1 |
| 2007 | ChemDB update - full-text search and virtual chemical spaceabstractUNLABELLED: ChemDB is a chemical database containing nearly 5M commercially available small molecules, important for use as synthetic building blocks, probes in systems biology and as leads for the discovery of drugs and other useful compounds. The data is publicly available over the web for download and for targeted searches using a variety of powerful methods. The chemical data includes predicted or experimentally determined physicochemical properties, such as 3D structure, melting temperature and solubility. Recent developments include optimization of chemical structure (and substructure) retrieval algorithms, enabling full database searches in less than a second. A text-based search engine allows efficient searching of compounds based on over 65M annotations from over 150 vendors. When searching for chemicals by name, fuzzy text matching capabilities yield productive results even when the correct spelling of a chemical name is unknown, taking advantage of both systematic and common names. Finally, built in reaction models enable searches through virtual chemical space, consisting of hypothetical products readily synthesizable from the building blocks in ChemDB. AVAILABILITY: ChemDB and Supplementary Materials are available at http://cdb.ics.uci.edu. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jonathan H. Chen, Erik Linstead, Sanjay Joshua Swamidass, Dennis Wang, Pierre Baldi |
Bioinform. | 3 |
| 2006 | Functional Census of Mutation Sequence Spaces: The Example of p53 Cancer Rescue MutantsabstractMany biomedical problems relate to mutant functional properties across a sequence space of interest, e.g., flu, cancer, and HIV. Detailed knowledge of mutant properties and function improves medical treatment and prevention. A functional census of p53 cancer rescue mutants would aid the search for cancer treatments from p53 mutant rescue. We devised a general methodology for conducting a functional census of a mutation sequence space by choosing informative mutants early. The methodology was tested in a double-blind predictive test on the functional rescue property of 71 novel putative p53 cancer rescue mutants iteratively predicted in sets of three (24 iterations). The first double-blind 15-point moving accuracy was 47 percent and the last was 86 percent; r = 0.01 before an epiphanic 16th iteration and r = 0.92 afterward. Useful mutants were chosen early (overall r = 0.80). Code and data are freely available (http://www.igb.uci.edu/research/research.html, corresponding authors: R.H.L. for computation and R.K.B. for biology). Samuel A. Danziger, Sanjay Joshua Swamidass, Jue Zeng, Lawrence R. Dearth, Jonathan H. Chen, Jianlin Cheng, Vinh P. Hoang, Hiroto Saigo, Ray Luo 0001, Pierre Baldi, Rainer K. Brachmann, Richard H. Lathrop |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2005 | ChemDB: a public database of small molecules and related chemoinformatics resourcesabstractMOTIVATION: The development of chemoinformatics has been hampered by the lack of large, publicly available, comprehensive repositories of molecules, in particular of small molecules. Small molecules play a fundamental role in organic chemistry and biology. They can be used as combinatorial building blocks for chemical synthesis, as molecular probes in chemical genomics and systems biology, and for the screening and discovery of new drugs and other useful compounds. RESULTS: We describe ChemDB, a public database of small molecules available on the Web. ChemDB is built using the digital catalogs of over a hundred vendors and other public sources and is annotated with information derived from these sources as well as from computational methods, such as predicted solubility and three-dimensional structure. It supports multiple molecular formats and is periodically updated, automatically whenever possible. The current version of the database contains approximately 4.1 million commercially available compounds and 8.2 million counting isomers. The database includes a user-friendly graphical interface, chemical reactions capabilities, as well as unique search capabilities. AVAILABILITY: Database and datasets are available on http://cdb.ics.uci.edu. Jonathan H. Chen, Sanjay Joshua Swamidass, Yimeng Dou, Jocelyne Bruand, Pierre Baldi |
Bioinform. | 2 |
| 2005 | Graph kernels for chemical informatics
Liva Ralaivola, Sanjay Joshua Swamidass, Hiroto Saigo, Pierre Baldi |
Neural Networks | 2 |