VLDB 2026 Research / reviewers in the wild / expert
Gordon B. Mills
dblp:66/3017
· DBLP profile ↗
16ranked-venue papers
0as first author
4since 2021 · last 2025
0000-0002-0144-9614ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Language of Stains: Tokenization Enhances Multiplex Immunofluorescence and Histology Image Synthesis
Zachary S. Sims, Sandhya Govindarajan, Gordon B. Mills, Ece Eksi, Young Hwan Chang |
MICCAI (4) | 3 |
| 2025 | COEXIST: Coordinated single-cell integration of serial multiplexed tissue imagesabstractMultiplexed tissue imaging (MTI) and other spatial profiling technologies commonly utilize serial tissue sectioning to comprehensively profile samples by imaging each section with unique biomarker panels or assays. The dependence on serial sections is attributed to technological limitations of MTI panel size or incompatible multi-assay protocols. Although image registration can align serially sectioned MTIs, integration at the single-cell level poses a challenge due to inherent biological heterogeneity. Existing computational methods overlook both cell population heterogeneity across modalities and spatial information, which are critical for effectively completing this task. To address this problem, we first use Monte-Carlo simulations to estimate the overlap between serial 5μm-thick sections. We then introduce COEXIST, a novel algorithm that synergistically combines shared molecular profiles with spatial information to seamlessly integrate serial sections at the single-cell level. We demonstrate COEXIST necessity and performance across several applications. These include combining MTI panels for improved spatial single-cell profiling, rectification of miscalled cell phenotypes using a single MTI panel, and the comparison of MTI platforms at single-cell resolution. COEXIST not only elevates MTI platform validation but also overcomes the constraints of MTI's panel size and the limitation of full nuclei on a single slide, capturing more intact nuclei in consecutive sections and thus enabling deeper profiling of cell lineages and functional states. Robert T. Heussner, Cameron F. Watson, Christopher Z. Eddy, Eric M. Cramer, Allison L. Creason, Gordon B. Mills, Young Hwan Chang |
PLoS Comput. Biol. | 7 |
| 2022 | RPPA SPACE: an R package for normalization and quantitation of Reverse-Phase Protein Array dataabstractSUMMARY: Reverse-Phase Protein Array (RPPA) is a robust high-throughput, cost-effective platform for quantitatively measuring proteins in biological specimens. However, converting raw RPPA data into normalized, analysis-ready data remains a challenging task. Here, we present the RPPA SPACE (RPPA Superposition Analysis and Concentration Evaluation) R package, a substantially improved successor to SuperCurve, to meet that challenge. SuperCurve has been used to normalize over 170 000 samples to date. RPPA SPACE allows exclusion of poor-quality samples from the normalization process to improve the quality of the remaining samples. It also features a novel quality-control metric, 'noise', that estimates the level of random errors present in each RPPA slide. The noise metric can help to determine the quality and reliability of the data. In addition, RPPA SPACE has simpler input requirements and is more flexible than SuperCurve, it is much faster with greatly improved error reporting. AVAILABILITY AND IMPLEMENTATION: The standalone RPPA SPACE R package, tutorials and sample data are available via https://rppa.space/, CRAN (https://cran.r-project.org/web/packages/RPPASPACE/index.html) and GitHub (https://github.com/MD-Anderson-Bioinformatics/RPPASPACE). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Huma Shehwana, Shwetha V. Kumar, James M. Melott, Mary A. Rohrdanz, Chris Wakefield, Zhenlin Ju, Doris R. Siwak, Yiling Lu, Bradley M. Broom, John N. Weinstein, Gordon B. Mills, Rehan Akbani |
Bioinform. | 11 |
| 2021 | mi-IsoNet: systems-scale microRNA landscape reveals rampant isoform-mediated gain of target interaction diversity and signaling specificityabstractMicroRNA (miRNA) is not a single sequence, but a series of multiple variants (also termed isomiRs) with sequence and expression heterogeneity. Whether and how these isoforms contribute to functional variation and complexity at the systems and network levels remain largely unknown. To explore this question systematically, we comprehensively analyzed the expression of small RNAs and their target sites to interrogate functional variations between novel isomiRs and their canonical miRNA sequences. Our analyses of the pan-cancer landscape of miRNA expression indicate that multiple isomiRs generated from the same miRNA locus often exhibit remarkable variation in their sequence, expression and function. We interrogated abundant and differentially expressed 5' isomiRs with novel seed sequences via seed shifting and identified many potential novel targets of these 5' isomiRs that would expand interaction capabilities between small RNAs and mRNAs, rewiring regulatory networks and increasing signaling circuit complexity. Further analyses revealed that some miRNA loci might generate diverse dominant isomiRs that often involved isomiRs with varied seeds and arm-switching, suggesting a selective advantage of multiple isomiRs in regulating gene expression. Finally, experimental validation indicated that isomiRs with shifted seed sequences could regulate novel target mRNAs and therefore contribute to regulatory network rewiring. Our analysis uncovers a widespread expansion of isomiR and mRNA interaction networks compared with those seen in canonical small RNA analysis; this expansion suggests global gene regulation network perturbations by alternative small RNA variants or isoforms. Taken together, the variations in isomiRs that occur during miRNA processing and maturation are likely to play a far more complex and plastic role in gene regulation than previously anticipated. Kara M. Cirillo, Robert A. Marick, Xing Yin, Xu Hua, Gordon B. Mills, Nidhi Sahni, S. Stephen Yi |
Briefings Bioinform. | 8 |
| 2017 | Molecular heterogeneity at the network level: high-dimensional testing, clustering and a TCGA case studyabstractMOTIVATION: Molecular pathways and networks play a key role in basic and disease biology. An emerging notion is that networks encoding patterns of molecular interplay may themselves differ between contexts, such as cell type, tissue or disease (sub)type. However, while statistical testing of differences in mean expression levels has been extensively studied, testing of network differences remains challenging. Furthermore, since network differences could provide important and biologically interpretable information to identify molecular subgroups, there is a need to consider the unsupervised task of learning subgroups and networks that define them. This is a nontrivial clustering problem, with neither subgroups nor subgroup-specific networks known at the outset. RESULTS: We leverage recent ideas from high-dimensional statistics for testing and clustering in the network biology setting. The methods we describe can be applied directly to most continuous molecular measurements and networks do not need to be specified beforehand. We illustrate the ideas and methods in a case study using protein data from The Cancer Genome Atlas (TCGA). This provides evidence that patterns of interplay between signalling proteins differ significantly between cancer types. Furthermore, we show how the proposed approaches can be used to learn subtypes and the molecular networks that define them. AVAILABILITY AND IMPLEMENTATION: As the Bioconductor package nethet. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Nicolas Städler, Frank Dondelinger, Steven M. Hill, Rehan Akbani, Yiling Lu, Gordon B. Mills, Sach Mukherjee |
Bioinform. | 6 |
| 2015 | Development of a robust classifier for quality control of reverse-phase protein arraysabstractMOTIVATION: High-throughput reverse-phase protein array (RPPA) technology allows for the parallel measurement of protein expression levels in approximately 1000 samples. However, the many steps required in the complex protocol (sample lysate preparation, slide printing, hybridization, washing and amplified detection) may create substantial variability in data quality. We are not aware of any other quality control algorithm that is tuned to the special characteristics of RPPAs. RESULTS: We have developed a novel classifier for quality control of RPPA experiments using a generalized linear model and logistic function. The outcome of the classifier, ranging from 0 to 1, is defined as the probability that a slide is of good quality. After training, we tested the classifier using two independent validation datasets. We conclude that the classifier can distinguish RPPA slides of good quality from those of poor quality sufficiently well such that normalization schemes, protein expression patterns and advanced biological analyses will not be drastically impacted by erroneous measurements or systematic variations. AVAILABILITY AND IMPLEMENTATION: The classifier, implemented in the "SuperCurve" R package, can be freely downloaded at http://bioinformatics.mdanderson.org/main/OOMPA:Overview or http://r-forge.r-project.org/projects/supercurve/. The data used to develop and validate the classifier are available at http://bioinformatics.mdanderson.org/MOAR. Zhenlin Ju, Paul L. Roebuck, Doris R. Siwak, Nianxiang Zhang, Yiling Lu, Michael A. Davies, Rehan Akbani, John N. Weinstein, Gordon B. Mills, Kevin R. Coombes |
Bioinform. | 10 |
| 2014 | Bias from removing read duplication in ultra-deep sequencing experimentsabstractMOTIVATION: Identifying subclonal mutations and their implications requires accurate estimation of mutant allele fractions from possibly duplicated sequencing reads. Removing duplicate reads assumes that polymerase chain reaction amplification from library constructions is the primary source. The alternative-sampling coincidence from DNA fragmentation-has not been systematically investigated. RESULTS: With sufficiently high-sequencing depth, sampling-induced read duplication is non-negligible, and removing duplicate reads can overcorrect read counts, causing systemic biases in variant allele fraction and copy number variation estimations. Minimal overcorrection occurs when duplicate reads are identified accounting for their mate reads, inserts are of a variety of lengths and samples are sequenced in separate batches. We investigate sampling-induced read duplication in deep sequencing data with 500× to 2000× duplicates-removed sequence coverage. We provide a quantitative solution to overcorrection and guidance for effective designs of deep sequencing platforms that facilitate accurate estimation of variant allele fraction and copy number variation. AVAILABILITY AND IMPLEMENTATION: A Python implementation is freely available at https://bitbucket.org/wanding/duprecover/overview CONTACT: : [email protected], [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Wanding Zhou, Tenghui Chen, Agda Karina Eterovic, Funda Meric-Bernstam, Gordon B. Mills, Ken Chen 0001 |
Bioinform. | 6 |
| 2013 | Network inference using steady-state data and Goldbeter-Koshland kineticsabstractMOTIVATION: Network inference approaches are widely used to shed light on regulatory interplay between molecular players such as genes and proteins. Biochemical processes underlying networks of interest (e.g. gene regulatory or protein signalling networks) are generally nonlinear. In many settings, knowledge is available concerning relevant chemical kinetics. However, existing network inference methods for continuous, steady-state data are typically rooted in statistical formulations, which do not exploit chemical kinetics to guide inference. RESULTS: Herein, we present an approach to network inference for steady-state data that is rooted in non-linear descriptions of biochemical mechanism. We use equilibrium analysis of chemical kinetics to obtain functional forms that are in turn used to infer networks using steady-state data. The approach we propose is directly applicable to conventional steady-state gene expression or proteomic data and does not require knowledge of either network topology or any kinetic parameters. We illustrate the approach in the context of protein phosphorylation networks, using data simulated from a recent mechanistic model and proteomic data from cancer cell lines. In the former, the true network is known and used for assessment, whereas in the latter, results are compared against known biochemistry. We find that the proposed methodology is more effective at estimating network topology than methods based on linear models. AVAILABILITY: mukherjeelab.nki.nl/CODE/GK_Kinetics.zip CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Chris J. Oates, Bryan T. J. Hennessy, Yiling Lu, Gordon B. Mills, Sach Mukherjee |
Bioinform. | 4 |
| 2013 | Perturbation Biology: Inferring Signaling Networks in Cellular SystemsabstractWe present a powerful experimental-computational technology for inferring network models that predict the response of cells to perturbations, and that may be useful in the design of combinatorial therapy against cancer. The experiments are systematic series of perturbations of cancer cell lines by targeted drugs, singly or in combination. The response to perturbation is quantified in terms of relative changes in the measured levels of proteins, phospho-proteins and cellular phenotypes such as viability. Computational network models are derived de novo, i.e., without prior knowledge of signaling pathways, and are based on simple non-linear differential equations. The prohibitively large solution space of all possible network models is explored efficiently using a probabilistic algorithm, Belief Propagation (BP), which is three orders of magnitude faster than standard Monte Carlo methods. Explicit executable models are derived for a set of perturbation experiments in SKMEL-133 melanoma cell lines, which are resistant to the therapeutically important inhibitor of RAF kinase. The resulting network models reproduce and extend known pathway biology. They empower potential discoveries of new molecular interactions and predict efficacious novel drug perturbations, such as the inhibition of PLK1, which is verified experimentally. This technology is suitable for application to larger systems in diverse areas of molecular biology. Evan J. Molinelli, Anil Korkut, Martin L. Miller, Nicholas Paul Gauthier, Xiaohong Jing, Poorvi Kaushik, Gordon B. Mills, David B. Solit, Christine A. Pratilas, Martin Weigt, Alfredo Braunstein, Andrea Pagnani, Riccardo Zecchina, Chris Sander |
PLoS Comput. Biol. | 9 |
| 2012 | Bayesian Inference of Signaling Network Topology in a Cancer Cell LineabstractMOTIVATION: Protein signaling networks play a key role in cellular function, and their dysregulation is central to many diseases, including cancer. To shed light on signaling network topology in specific contexts, such as cancer, requires interrogation of multiple proteins through time and statistical approaches to make inferences regarding network structure. RESULTS: In this study, we use dynamic Bayesian networks to make inferences regarding network structure and thereby generate testable hypotheses. We incorporate existing biology using informative network priors, weighted objectively by an empirical Bayes approach, and exploit a connection between variable selection and network inference to enable exact calculation of posterior probabilities of interest. The approach is computationally efficient and essentially free of user-set tuning parameters. Results on data where the true, underlying network is known place the approach favorably relative to existing approaches. We apply these methods to reverse-phase protein array time-course data from a breast cancer cell line (MDA-MB-468) to predict signaling links that we independently validate using targeted inhibition. The methods proposed offer a general approach by which to elucidate molecular networks specific to biological context, including, but not limited to, human cancers. AVAILABILITY: http://mukherjeelab.nki.nl/DBN (code and data). Steven M. Hill, Yiling Lu, Jennifer Molina, Laura Heiser, Paul T. Spellman, Terence P. Speed, Joe W. Gray, Gordon B. Mills, Sach Mukherjee |
Bioinform. | 8 |
| 2012 | Network inference using steady-state data and Goldbeter-koshland kineticsabstractAbstract Motivation: Network inference approaches are widely used to shed light on regulatory interplay between molecular players such as genes and proteins. Biochemical processes underlying networks of interest (e.g. gene regulatory or protein signalling networks) are generally nonlinear. In many settings, knowledge is available concerning relevant chemical kinetics. However, existing network inference methods for continuous, steady-state data are typically rooted in statistical formulations, which do not exploit chemical kinetics to guide inference. Results: Herein, we present an approach to network inference for steady-state data that is rooted in non-linear descriptions of biochemical mechanism. We use equilibrium analysis of chemical kinetics to obtain functional forms that are in turn used to infer networks using steady-state data. The approach we propose is directly applicable to conventional steady-state gene expression or proteomic data and does not require knowledge of either network topology or any kinetic parameters. We illustrate the approach in the context of protein phosphorylation networks, using data simulated from a recent mechanistic model and proteomic data from cancer cell lines. In the former, the true network is known and used for assessment, whereas in the latter, results are compared against known biochemistry. We find that the proposed methodology is more effective at estimating network topology than methods based on linear models. Availability: mukherjeelab.nki.nl/CODE/GK_Kinetics.zip Contact: [email protected]; [email protected] Supplementary Information: Supplementary data are available at Bioinformatics online. Chris J. Oates, Bryan T. J. Hennessy, Yiling Lu, Gordon B. Mills, Sach Mukherjee |
Bioinform. | 4 |
| 2010 | Exposing the cancer genome atlas as a SPARQL endpoint
Helena F. Deus, Diogo F. Veiga, Pablo R. Freire, John N. Weinstein, Gordon B. Mills, Jonas S. Almeida |
J. Biomed. Informatics | 5 |
| 2009 | Serial dilution curve: a new method for analysis of reverse phase protein array dataabstractUNLABELLED: Reverse phase protein arrays (RPPAs) are a powerful high-throughput tool for measuring protein concentrations in a large number of samples. In RPPA technology, the original samples are often diluted successively multiple times, forming dilution series to extend the dynamic range of the measurements and to increase confidence in quantitation. An RPPA experiment is equivalent to running multiple ELISA assays concurrently except that there is usually no known protein concentration from which one can construct a standard response curve. Here, we describe a new method called 'serial dilution curve for RPPA data analysis'. Compared with the existing methods, the new method has the advantage of using fewer parameters and offering a simple way of visualizing the raw data. We showed how the method can be used to examine data quality and to obtain robust quantification of protein concentrations. AVAILABILITY: A computer program in R for using serial dilution curve for RPPA data analysis is freely available at http://odin.mdacc.tmc.edu/~zhangli/RPPA. Qingyi Wei, Gordon B. Mills, Kevin R. Coombes |
Bioinform. | 5 |
| 2008 | Bayesian models based on test statistics for multiple hypothesis testing problemsabstractMOTIVATION: We propose a Bayesian method for the problem of multiple hypothesis testing that is routinely encountered in bioinformatics research, such as the differential gene expression analysis. Our algorithm is based on modeling the distributions of test statistics under both null and alternative hypotheses. We substantially reduce the complexity of the process of defining posterior model probabilities by modeling the test statistics directly instead of modeling the full data. Computationally, we apply a Bayesian FDR approach to control the number of rejections of null hypotheses. To check if our model assumptions for the test statistics are valid for various bioinformatics experiments, we also propose a simple graphical model-assessment tool. RESULTS: Using extensive simulations, we demonstrate the performance of our models and the utility of the model-assessment tool. In the end, we apply the proposed methodology to an siRNA screening and a gene expression experiment. Yiling Lu, Gordon B. Mills |
Bioinform. | 3 |
| 2008 | RPPAML/RIMS: A metadata format and an information management system for reverse phase protein arraysabstractBACKGROUND: Reverse Phase Protein Arrays (RPPA) are convenient assay platforms to investigate the presence of biomarkers in tissue lysates. As with other high-throughput technologies, substantial amounts of analytical data are generated. Over 1,000 samples may be printed on a single nitrocellulose slide. Up to 100 different proteins may be assessed using immunoperoxidase or immunoflorescence techniques in order to determine relative amounts of protein expression in the samples of interest. RESULTS: In this report an RPPA Information Management System (RIMS) is described and made available with open source software. In order to implement the proposed system, we propose a metadata format known as reverse phase protein array markup language (RPPAML). RPPAML would enable researchers to describe, document and disseminate RPPA data. The complexity of the data structure needed to describe the results and the graphic tools necessary to visualize them require a software deployment distributed between a client and a server application. This was achieved without sacrificing interoperability between individual deployments through the use of an open source semantic database, S3DB. This data service backbone is available to multiple client side applications that can also access other server side deployments. The RIMS platform was designed to interoperate with other data analysis and data visualization tools such as Cytoscape. CONCLUSION: The proposed RPPAML data format hopes to standardize RPPA data. Standardization of data would result in diverse client applications being able to operate on the same set of data. Additionally, having data in a standard format would enable data dissemination and data analysis. Romesh Stanislaus, Mark Carey, Helena F. Deus, Kevin R. Coombes, Bryan T. J. Hennessy, Gordon B. Mills, Jonas S. Almeida |
BMC Bioinform. | 6 |
| 2007 | Non-parametric quantification of protein lysate arraysabstractMOTIVATION: Proteins play a crucial role in biological activity, so much can be learned from measuring protein expression and post-translational modification quantitatively. The reverse-phase protein lysate arrays allow us to quantify the relative expression levels of a protein in many different cellular samples simultaneously. Existing approaches to quantify protein arrays use parametric response curves fit to dilution series data. The results can be biased when the parametric function does not fit the data. RESULTS: We propose a non-parametric approach which adapts to any monotone response curve. The non-parametric approach is shown to be promising via both simulation and real data studies; it reduces the bias due to model misspecification and protects against outliers in the data. The non-parametric approach enables more reliable quantification of protein lysate arrays. AVAILABILITY: Code to implement the proposed method in the statistical package R is available at: http://odin.mdacc.tmc.edu/jhu/lysatearray-analysis/ Xuming He 0002, Keith A. Baggerly, Kevin R. Coombes, Bryan T. J. Hennessy, Gordon B. Mills |
Bioinform. | 6 |