VLDB 2026 Research / reviewers in the wild / expert
Andrew I. Su
dblp:62/2725
· DBLP profile ↗
15ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0002-9859-4104ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A case-based explainable graph neural network framework for mechanistic drug repositioningabstractDrug repositioning offers a cost-effective alternative to traditional drug development by identifying new uses for existing drugs. Recent advances leverage Graph Neural Networks (GNNs) to model complex biological data, showing promise in predicting novel drug-disease associations; however, these frameworks often lack explainability, a critical factor for validating predictions and understanding drug mechanisms. Here, we introduce Drug-Based Reasoning Explainer (DBR-X), an explainable GNN model that integrates a link-prediction module with a path-identification module to generate interpretable and faithful explanations. When benchmarked against other GNN-based link-prediction frameworks, DBR-X achieves superior performance in identifying known drug-disease associations, demonstrating higher accuracy across all evaluation metrics. The quality of DBR-X biological explanations was evaluated through multiple complementary approaches, including comparison with manually curated drug mechanisms, assessment of explanation faithfulness using deletion and insertion studies, and measurement of stability under graph perturbations. Together, these results show that DBR-X advances the state of the art in drug repositioning while providing multi-hop mechanistic explanations that can facilitate the translation of computational predictions into clinical applications. Availability and implementation: DBR-X package is freely accessible from online repository https://github.com/SuLab/DBR-X. Adriana Carolina Gonzalez-Cavazos, Roger Tu, Meghamala Sinha, Andrew I. Su |
Bioinform. | 4 |
| 2023 | BioThings Explorer: a query engine for a federated knowledge graph of biomedical APIsabstractKnowledge graphs are an increasingly common data structure for representing biomedical information. These knowledge graphs can easily represent heterogeneous types of information, and many algorithms and tools exist for querying and analyzing graphs. Biomedical knowledge graphs have been used in a variety of applications, including drug repurposing, identification of drug targets, prediction of drug side effects, and clinical decision support. Typically, knowledge graphs are constructed by centralization and integration of data from multiple disparate sources. Here, we describe BioThings Explorer, an application that can query a virtual, federated knowledge graph derived from the aggregated information in a network of biomedical web services. BioThings Explorer leverages semantically precise annotations of the inputs and outputs for each resource, and automates the chaining of web service calls to execute multi-step graph queries. Because there is no large, centralized knowledge graph to maintain, BioThing Explorer is distributed as a lightweight application that dynamically retrieves information at query time. More information can be found at https://explorer.biothings.io, and code is available at https://github.com/biothings/biothings_explorer. Jackson Callaghan, Colleen H. Xu, Jiwen Xin, Marco Alvarado Cano, Anders Riutta, Eric Zhou, Rohan Juneja, Yao Yao 0007, Madhumita Narayan, Kristina Hanspers, Ayushi Agrawal, Alexander R. Pico, Chunlei Wu, Andrew I. Su |
Bioinform. | 14 |
| 2023 | Schema Playground: a tool for authoring, extending, and using metadata schemas to improve FAIRness of biomedical dataabstractBACKGROUND: Biomedical researchers are strongly encouraged to make their research outputs more Findable, Accessible, Interoperable, and Reusable (FAIR). While many biomedical research outputs are more readily accessible through open data efforts, finding relevant outputs remains a significant challenge. Schema.org is a metadata vocabulary standardization project that enables web content creators to make their content more FAIR. Leveraging Schema.org could benefit biomedical research resource providers, but it can be challenging to apply Schema.org standards to biomedical research outputs. We created an online browser-based tool that empowers researchers and repository developers to utilize Schema.org or other biomedical schema projects. RESULTS: Our browser-based tool includes features which can help address many of the barriers towards Schema.org-compliance such as: The ability to easily browse for relevant Schema.org classes, the ability to extend and customize a class to be more suitable for biomedical research outputs, the ability to create data validation to ensure adherence of a research output to a customized class, and the ability to register a custom class to our schema registry enabling others to search and re-use it. We demonstrate the use of our tool with the creation of the Outbreak.info schema-a large multi-class schema for harmonizing various COVID-19 related resources. CONCLUSIONS: We have created a browser-based tool to empower biomedical research resource providers to leverage Schema.org classes to make their research outputs more FAIR. Marco Alvarado Cano, Ginger Tsueng, Xinghua Zhou, Jiwen Xin, Laura D. Hughes, Julia Mullen, Andrew I. Su, Chunlei Wu |
BMC Bioinform. | 7 |
| 2022 | BioThings SDK: a toolkit for building high-performance data APIs in biomedical researchabstractSUMMARY: To meet the increased need of making biomedical resources more accessible and reusable, Web Application Programming Interfaces (APIs) or web services have become a common way to disseminate knowledge sources. The BioThings APIs are a collection of high-performance, scalable, annotation as a service APIs that automate the integration of biological annotations from disparate data sources. This collection of APIs currently includes MyGene.info, MyVariant.info and MyChem.info for integrating annotations on genes, variants and chemical compounds, respectively. These APIs are used by both individual researchers and application developers to simplify the process of annotation retrieval and identifier mapping. Here, we describe the BioThings Software Development Kit (SDK), a generalizable and reusable toolkit for integrating data from multiple disparate data sources and creating high-performance APIs. This toolkit allows users to easily create their own BioThings APIs for any data type of interest to them, as well as keep APIs up-to-date with their underlying data sources. AVAILABILITY AND IMPLEMENTATION: The BioThings SDK is built in Python and released via PyPI (https://pypi.org/project/biothings/). Its source code is hosted at its github repository (https://github.com/biothings/biothings.api). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Sebastien Lelong, Xinghua Zhou, Cyrus Afrasiabi, Zhongchao Qian, Marco Alvarado Cano, Ginger Tsueng, Jiwen Xin, Julia Mullen, Yao Yao 0007, Ricardo Avila, Greg Taylor, Andrew I. Su, Chunlei Wu |
Bioinform. | 12 |
| 2022 | Design and application of a knowledge network for automatic prioritization of drug mechanismsabstractMOTIVATION: Drug repositioning is an attractive alternative to de novo drug discovery due to reduced time and costs to bring drugs to market. Computational repositioning methods, particularly non-black-box methods that can account for and predict a drug's mechanism, may provide great benefit for directing future development. By tuning both data and algorithm to utilize relationships important to drug mechanisms, a computational repositioning algorithm can be trained to both predict and explain mechanistically novel indications. RESULTS: In this work, we examined the 123 curated drug mechanism paths found in the drug mechanism database (DrugMechDB) and after identifying the most important relationships, we integrated 18 data sources to produce a heterogeneous knowledge graph, MechRepoNet, capable of capturing the information in these paths. We applied the Rephetio repurposing algorithm to MechRepoNet using only a subset of relationships known to be mechanistic in nature and found adequate predictive ability on an evaluation set with AUROC value of 0.83. The resulting repurposing model allowed us to prioritize paths in our knowledge graph to produce a predicted treatment mechanism. We found that DrugMechDB paths, when present in the network were rated highly among predicted mechanisms. We then demonstrated MechRepoNet's ability to use mechanistic insight to identify a drug's mechanistic target, with a mean reciprocal rank of 0.525 on a test set of known drug-target interactions. Finally, we walked through repurposing examples of the anti-cancer drug imatinib for use in the treatment of asthma, and metolazone for use in the treatment of osteoporosis, to demonstrate this method's utility in providing mechanistic insight into repurposing predictions it provides. AVAILABILITY AND IMPLEMENTATION: The Python code to reproduce the entirety of this analysis is available at: https://github.com/SuLab/MechRepoNet (archived at https://doi.org/10.5281/zenodo.6456335). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Michael Mayers, Roger Tu, Dylan Steinecke, Tong Shu Li, Núria Queralt-Rosinach, Andrew I. Su |
Bioinform. | 6 |
| 2020 | Applying citizen science to gene, drug and disease relationship extraction from biomedical abstractsabstractMOTIVATION: Biomedical literature is growing at a rate that outpaces our ability to harness the knowledge contained therein. To mine valuable inferences from the large volume of literature, many researchers use information extraction algorithms to harvest information in biomedical texts. Information extraction is usually accomplished via a combination of manual expert curation and computational methods. Advances in computational methods usually depend on the time-consuming generation of gold standards by a limited number of expert curators. Citizen science is public participation in scientific research. We previously found that citizen scientists are willing and capable of performing named entity recognition of disease mentions in biomedical abstracts, but did not know if this was true with relationship extraction (RE). RESULTS: In this article, we introduce the Relationship Extraction Module of the web-based application Mark2Cure (M2C) and demonstrate that citizen scientists can perform RE. We confirm the importance of accurate named entity recognition on user performance of RE and identify design issues that impacted data quality. We find that the data generated by citizen scientists can be used to identify relationship types not currently available in the M2C Relationship Extraction Module. We compare the citizen science-generated data with algorithm-mined data and identify ways in which the two approaches may complement one another. We also discuss opportunities for future improvement of this system, as well as the potential synergies between citizen science, manual biocuration and natural language processing. AVAILABILITY AND IMPLEMENTATION: Mark2Cure platform: https://mark2cure.org; Mark2Cure source code: https://github.com/sulab/mark2cure; and data and analysis code for this article: https://github.com/gtsueng/M2C_rel_nb. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ginger Tsueng, Max Nanis, Jennifer T. Fouquier, Michael Mayers, Benjamin M. Good, Andrew I. Su |
Bioinform. | 6 |
| 2019 | Time-resolved evaluation of compound repositioning predictions on a text-mined knowledge networkabstractBACKGROUND: Computational compound repositioning has the potential for identifying new uses for existing drugs, and new algorithms and data source aggregation strategies provide ever-improving results via in silico metrics. However, even with these advances, the number of compounds successfully repositioned via computational screening remains low. New strategies for algorithm evaluation that more accurately reflect the repositioning potential of a compound could provide a better target for future optimizations. RESULTS: Using a text-mined database, we applied a previously described network-based computational repositioning algorithm, yielding strong results via cross-validation, averaging 0.95 AUROC on test-set indications. However, to better approximate a real-world scenario, we built a time-resolved evaluation framework. At various time points, we built networks corresponding to prior knowledge for use as a training set, and then predicted on a test set comprised of indications that were subsequently described. This framework showed a marked reduction in performance, peaking in performance metrics with the 1985 network at an AUROC of .797. Examining performance reductions due to removal of specific types of relationships highlighted the importance of drug-drug and disease-disease similarity metrics. Using data from future timepoints, we demonstrate that further acquisition of these kinds of data may help improve computational results. CONCLUSIONS: Evaluating a repositioning algorithm using indications unknown to input network better tunes its ability to find emerging drug indications, rather than finding those which have been randomly withheld. Focusing efforts on improving algorithmic performance in a time-resolved paradigm may further improve computational repositioning predictions. Michael Mayers, Tong Shu Li, Núria Queralt-Rosinach, Andrew I. Su |
BMC Bioinform. | 4 |
| 2018 | Cross-linking BioThings APIs through JSON-LD to facilitate knowledge explorationabstractBACKGROUND: Application Programming Interfaces (APIs) are now widely used to distribute biological data. And many popular biological APIs developed by many different research teams have adopted Javascript Object Notation (JSON) as their primary data format. While usage of a common data format offers significant advantages, that alone is not sufficient for rich integrative queries across APIs. RESULTS: Here, we have implemented JSON for Linking Data (JSON-LD) technology on the BioThings APIs that we have developed, MyGene.info , MyVariant.info and MyChem.info . JSON-LD provides a standard way to add semantic context to the existing JSON data structure, for the purpose of enhancing the interoperability between APIs. We demonstrated several use cases that were facilitated by semantic annotations using JSON-LD, including simpler and more precise query capabilities as well as API cross-linking. CONCLUSIONS: We believe that this pattern offers a generalizable solution for interoperability of APIs in the life sciences. Jiwen Xin, Cyrus Afrasiabi, Sebastien Lelong, Julee Adesara, Ginger Tsueng, Andrew I. Su, Chunlei Wu |
BMC Bioinform. | 6 |
| 2016 | Crowdsourcing in biomedicine: challenges and opportunitiesabstractThe use of crowdsourcing to solve important but complex problems in biomedical and clinical sciences is growing and encompasses a wide variety of approaches. The crowd is diverse and includes online marketplace workers, health information seekers, science enthusiasts and domain experts. In this article, we review and highlight recent studies that use crowdsourcing to advance biomedicine. We classify these studies into two broad categories: (i) mining big data generated from a crowd (e.g. search logs) and (ii) active crowdsourcing via specific technical platforms, e.g. labor markets, wikis, scientific games and community challenges. Through describing each study in detail, we demonstrate the applicability of different methods in a variety of domains in biomedical research, including genomics, biocuration and clinical research. Furthermore, we discuss and highlight the strengths and limitations of different crowdsourcing platforms. Finally, we identify important emerging trends, opportunities and remaining challenges for future crowdsourcing research in biomedicine. Ritu Khare, Benjamin M. Good, Robert Leaman, Andrew I. Su, Zhiyong Lu |
Briefings Bioinform. | 4 |
| 2016 | Branch: an interactive, web-based tool for testing hypotheses and developing predictive modelsabstractUNLABELLED: Branch is a web application that provides users with the ability to interact directly with large biomedical datasets. The interaction is mediated through a collaborative graphical user interface for building and evaluating decision trees. These trees can be used to compose and test sophisticated hypotheses and to develop predictive models. Decision trees are built and evaluated based on a library of imported datasets and can be stored in a collective area for sharing and re-use. AVAILABILITY AND IMPLEMENTATION: Branch is hosted at http://biobranch.org/ and the open source code is available at http://bitbucket.org/sulab/biobranch/ CONTACTS: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Karthik Gangavarapu, Vyshakh Babji, Tobias Meißner, Andrew I. Su, Benjamin M. Good |
Bioinform. | 4 |
| 2015 | Omics Pipe: a community-based framework for reproducible multi-omics data analysisabstractMOTIVATION: Omics Pipe (http://sulab.scripps.edu/omicspipe) is a computational framework that automates multi-omics data analysis pipelines on high performance compute clusters and in the cloud. It supports best practice published pipelines for RNA-seq, miRNA-seq, Exome-seq, Whole-Genome sequencing, ChIP-seq analyses and automatic processing of data from The Cancer Genome Atlas (TCGA). Omics Pipe provides researchers with a tool for reproducible, open source and extensible next generation sequencing analysis. The goal of Omics Pipe is to democratize next-generation sequencing analysis by dramatically increasing the accessibility and reproducibility of best practice computational pipelines, which will enable researchers to generate biologically meaningful and interpretable results. RESULTS: Using Omics Pipe, we analyzed 100 TCGA breast invasive carcinoma paired tumor-normal datasets based on the latest UCSC hg19 RefSeq annotation. Omics Pipe automatically downloaded and processed the desired TCGA samples on a high throughput compute cluster to produce a results report for each sample. We aggregated the individual sample results and compared them to the analysis in the original publications. This comparison revealed high overlap between the analyses, as well as novel findings due to the use of updated annotations and methods. AVAILABILITY AND IMPLEMENTATION: Source code for Omics Pipe is freely available on the web (https://bitbucket.org/sulab/omics_pipe). Omics Pipe is distributed as a standalone Python package for installation (https://pypi.python.org/pypi/omics_pipe) and as an Amazon Machine Image in Amazon Web Services Elastic Compute Cloud that contains all necessary third-party software dependencies and databases (https://pythonhosted.org/omics_pipe/AWS_installation.html). Kathleen M. Fisch, Tobias Meißner, Louis Gioia, Jean-Christophe Ducom, Tristan M. Carland, Salvatore Loguercio, Andrew I. Su |
Bioinform. | 7 |
| 2013 | Crowdsourcing for bioinformaticsabstractMOTIVATION: Bioinformatics is faced with a variety of problems that require human involvement. Tasks like genome annotation, image analysis, knowledge-base population and protein structure determination all benefit from human input. In some cases, people are needed in vast quantities, whereas in others, we need just a few with rare abilities. Crowdsourcing encompasses an emerging collection of approaches for harnessing such distributed human intelligence. Recently, the bioinformatics community has begun to apply crowdsourcing in a variety of contexts, yet few resources are available that describe how these human-powered systems work and how to use them effectively in scientific domains. RESULTS: Here, we provide a framework for understanding and applying several different types of crowdsourcing. The framework considers two broad classes: systems for solving large-volume 'microtasks' and systems for solving high-difficulty 'megatasks'. Within these classes, we discuss system types, including volunteer labor, games with a purpose, microtask markets and open innovation contests. We illustrate each system type with successful examples in bioinformatics and conclude with a guide for matching problems to crowdsourcing solutions that highlights the positives and negatives of different approaches. Benjamin M. Good, Andrew I. Su |
Bioinform. | 2 |
| 2013 | Ten Simple Rules for Cultivating Open Science and Collaborative R&DabstractHow can we address the complexity and cost of applying science to societal challenges?
Open science and collaborative R&D may help [1]–[3]. Open science has been described as “a research accelerator” [4]. Open science implies open access [5] but goes beyond it: “Imagine a connected online web of scientific knowledge that integrates and connects data, computer code, chains of scientific reasoning, descriptions of open problems, and beyond …. tightly integrated with a scientific social web that directs scientists' attention where it is most valuable, releasing enormous collaborative potential.” [1].
Open science and collaborative approaches are often described as open source, by analogy with open-source software such as the operating system Linux which powers Google and Amazon—collaboratively created software which is free to use and adapt, and popular for Internet infrastructure and scientific research [6], [7]. However, this use of “open source” is unclear. Some people use “open source” when a project's results are free to use, others when a project's process is highly collaborative [4].
It is clearer to classify open source and open science within a broader class of collaborative R&D, which can be defined as scalable collaboration (usually enabled by information technology) across organizational boundaries to solve R&D challenges [8].
Many approaches to open science and collaborative R&D have been tried [1], [9]. The Gene Wiki has created over 10,000 Wikipedia articles, and aims to provide one for every notable human gene [10]. The crowdsourcing platform InnoCentive has reportedly facilitated solutions to roughly half of the thousands of technical problems posed on the site, including many in life sciences such as the $1 million ALS Biomarker Prize [11]. Other examples include prizes (X-Prize [12]), scientific games (FoldIt [13]), and licensing schemes inspired by open-source software (BIOS [14]).
Collaborative R&D approaches vary in openness [15]. In some approaches, the R&D process and outputs are open to all—for example, open-science projects like the Gene Wiki described above. In other approaches which demonstrate what might be called controlled collaboration, there are strong controls on who contributes and benefits—for example, computational platforms like Collaborative Drug Discovery or InnoCentive that support both commercial and nonprofit research [9], [11].
Collaborative approaches can unleash innovation from unforeseen sources, as with crowdsourcing health technologies [11]–[13], [16]. They may help in global challenges like drug development [17], as with India's OSDD (Open Source Drug Discovery) project that recruited over 7,000 volunteers [16] and an open-source drug synthesis project that improved an existing drug without increasing its cost [18].
If you want to apply open science and collaborative R&D, what principles are useful? We suggest Ten Simple Rules for Cultivating Open Science and Collaborative R&D. We also offer eight conversational interviews exploring life experiences that led to these rules (Box 1).
Box 1. Conversations on Open Science and Collaborative R&D
Many commentators have considered challenges in translating open science and collaborative methods to biomedical research [2]–[4], [9], [17], [20], [24], [26], [28], [29]. How can protecting intellectual property be balanced with freeing researchers to build on previous knowledge? If R&D results are collaboratively created and freely available, who will take responsibility for costly clinical trials and quality control? What will be the Linux of open-source R&D?
To explore such challenges and convey life experiences in biomedical open science and collaborative R&D, we offer eight conversational interviews by the first author of this article as supplementary material. The conversations were done on behalf of the Results for Development Institute and are with:
Alph Bingham, cofounder of InnoCentive (Text S1)
Barry Bunin, CEO of Collaborative Drug Discovery (Text S2)
Leslie Chan, open access pioneer and director of Bioline International (Text S3)
Aled Edwards, director of the Structural Genomics Consortium (Text S4)
Benjamin Good, coleader of the Gene Wiki initiative (Text S5)
Bernard Munos, pharmaceutical innovation thought leader (Text S6)
Zakir Thomas, director of India's Open Source Drug Discovery (OSDD) project (Text S7)
Matt Todd, open science and drug development pioneer (Text S8) Hassan Masum, Aarthi Rao, Benjamin M. Good, Matthew H. Todd, Aled M. Edwards, Leslie Chan, Barry A. Bunin, Andrew I. Su, Zakir Thomas, Philip E. Bourne |
PLoS Comput. Biol. | 8 |
| 2011 | TCLUST: A Fast Method for Clustering Genome-Scale Expression DataabstractGenes with a common function are often hypothesized to have correlated expression levels in mRNA expression data, motivating the development of clustering algorithms for gene expression data sets. We observe that existing approaches do not scale well for large data sets, and indeed did not converge for the data set considered here. We present a novel clustering method TCLUST that exploits coconnectedness to efficiently cluster large, sparse expression data. We compare our approach with two existing clustering methods CAST and K-means which have been previously applied to clustering of gene-expression data with good performance results. Using a number of metrics, TCLUST is shown to be superior to or at least competitive with the other methods, while being much faster. We have applied this clustering algorithm to a genome-scale gene-expression data set and used gene set enrichment analysis to discover highly significant biological clusters. (Source code for TCLUST is downloadable at http://www.cse.ucsd.edu/~bdost/tclust.) Banu Dost, Chunlei Wu, Andrew I. Su, Vineet Bafna |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2006 | Comparative analysis of haplotype association mapping algorithmsabstractBACKGROUND: Finding the genetic causes of quantitative traits is a complex and difficult task. Classical methods for mapping quantitative trail loci (QTL) in miceuse an F2 cross between two strains with substantially different phenotype and an interval mapping method to compute confidence intervals at each position in the genome. This process requires significant resources for breeding and genotyping, and the data generated are usually only applicable to one phenotype of interest. Recently, we reported the application of a haplotype association mapping method which utilizes dense genotyping data across a diverse panel of inbred mouse strains and a marker association algorithm that is independent of any specific phenotype. As the availability of genotyping data grows in size and density, analysis of these haplotype association mapping methods should be of increasing value to the statistical genetics community. RESULTS: We describe a detailed comparative analysis of variations on our marker association method. In particular, we describe the use of inferred haplotypes from adjacent SNPs, parametric and nonparametric statistics, and control of multiple testing error. These results show that nonparametric methods are slightly better in the test cases we study, although the choice of test statistic may often be dependent on the specific phenotype and haplotype structure being studied. The use of multi-SNP windows to infer local haplotype structure is critical to the use of a diverse panel of inbred strains for QTL mapping. Finally, because the marginal effect of any single gene in a complex disease is often relatively small, these methods require the use of sensitive methods for controlling family-wise error. We also report our initial application of this method to phenotypes cataloged in the Mouse Phenome Database. CONCLUSION: The use of inbred strains of mice for QTL mapping has many advantages over traditional methods. However, there are also limitations in comparison to the traditional linkage analysis from F2 and RI lines. Application of these methods requires careful consideration of algorithmic choices based on both theoretical and practical factors. Our findings suggest general guidelines, though a complete evaluation of these methods can only be performed as more genetic data in complex diseases becomes available. Phillip McClurg, Mathew T. Pletcher, Tim Wiltshire, Andrew I. Su |
BMC Bioinform. | 4 |