VLDB 2026 Research / reviewers in the wild / expert
Michael Inouye
dblp:22/1580
· DBLP profile ↗
9ranked-venue papers
0as first author
4since 2021 · last 2024
0000-0001-9413-6520ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
4 papers |
Bioinformatics and computational biology · 82% Computational science and engineering · 18% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Emerging computing paradigms · 70% Reconfigurable computing and FPGAs · 30% |
Topics — the 14 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computational science and engineering
benchmark framework |
0.8 | 1 | 2024 | Pitfalls of machine learning models for protein-protein interaction networks · Bioinform. 2024 |
Bioinformatics and computational biology
functional genomics |
0.8 | 1 | 2024 | Pitfalls of machine learning models for protein-protein interaction networks · Bioinform. 2024 |
Bioinformatics and computational biology › protein analysis › protein-protein interaction › protein-protein interaction network analysis
protein-protein interaction network inference |
0.8 | 1 | 2024 | Pitfalls of machine learning models for protein-protein interaction networks · Bioinform. 2024 |
Bioinformatics and computational biology
protein-protein interaction prediction |
0.8 | 1 | 2024 | Pitfalls of machine learning models for protein-protein interaction networks · Bioinform. 2024 |
Reconfigurable computing and FPGAs
FPGA-based emulation |
0.5 | 1 | 2021 | RANC: Reconfigurable Architecture for Neuromorphic Computing · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2021 |
Emerging computing paradigms
neuromorphic computing |
0.5 | 1 | 2021 | RANC: Reconfigurable Architecture for Neuromorphic Computing · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2021 |
Emerging computing paradigms › neuromorphic computing
spiking neural network |
0.5 | 1 | 2021 | RANC: Reconfigurable Architecture for Neuromorphic Computing · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2021 |
Bioinformatics and computational biology › computational microbiology
microbiome analysis |
0.4 | 1 | 2019 | FastSpar: rapid and scalable correlation estimation for compositional data · Bioinform. 2019 |
Bioinformatics and computational biology › population genetics
ancestry inference |
0.3 | 1 | 2017 | FlashPCA2: principal component analysis of Biobank-scale genotype datasets · Bioinform. 2017 |
Bioinformatics and computational biology
population genetics |
0.3 | 1 | 2017 | FlashPCA2: principal component analysis of Biobank-scale genotype datasets · Bioinform. 2017 |
Emerging computing paradigms
neuromorphic hardware |
0.1 | 1 | 2021 | RANC: Reconfigurable Architecture for Neuromorphic Computing · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2021 |
Algorithms and data structures
parallel algorithms |
0.1 | 1 | 2019 | FastSpar: rapid and scalable correlation estimation for compositional data · Bioinform. 2019 |
Bioinformatics and computational biology › statistical genetics
genotype calling |
0.1 | 1 | 2007 | A genotype calling algorithm for the Illumina BeadArray platform · Bioinform. 2007 |
Bioinformatics and computational biology › genomics
genotyping |
0.1 | 1 | 2007 | A genotype calling algorithm for the Illumina BeadArray platform · Bioinform. 2007 |
Methods — techniques the papers use, named apart from their topics
unbiased estimator · 0.8reproducibility analysis · 0.8parallelization · 0.8machine learning · 0.8benchmarking · 0.8SparCC · 0.8c++ simulation · 0.5FPGA emulation · 0.5randomized algorithm · 0.3partial singular value decomposition · 0.3quality metric · 0.1perturbation analysis · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Pitfalls of machine learning models for protein-protein interaction networksabstractMOTIVATION: Protein-protein interactions (PPIs) are essential to understanding biological pathways as well as their roles in development and disease. Computational tools, based on classic machine learning, have been successful at predicting PPIs in silico, but the lack of consistent and reliable frameworks for this task has led to network models that are difficult to compare and discrepancies between algorithms that remain unexplained. RESULTS: To better understand the underlying inference mechanisms that underpin these models, we designed an open-source framework for benchmarking that accounts for a range of biological and statistical pitfalls while facilitating reproducibility. We use it to shed light on the impact of network topology and how different algorithms deal with highly connected proteins. By studying functional genomics-based and sequence-based models on human PPIs, we show their complementarity as the former performs best on lone proteins while the latter specializes in interactions involving hubs. We also show that algorithm design has little impact on performance with functional genomic data. We replicate our results between both human and S. cerevisiae data and demonstrate that models using functional genomics are better suited to PPI prediction across species. With rapidly increasing amounts of sequence and functional genomics data, our study provides a principled foundation for future construction, comparison, and application of PPI networks. AVAILABILITY AND IMPLEMENTATION: The code and data are available on GitHub: https://github.com/Llannelongue/B4PPI. Loïc Lannelongue, Michael Inouye |
Bioinform. | 2 |
| 2022 | Known allosteric proteins have central roles in genetic diseaseabstractAllostery is a form of protein regulation, where ligands that bind sites located apart from the active site can modify the activity of the protein. The molecular mechanisms of allostery have been extensively studied, because allosteric sites are less conserved than active sites, and drugs targeting them are more specific than drugs binding the active sites. Here we quantify the importance of allostery in genetic disease. We show that 1) known allosteric proteins are central in disease networks, contribute to genetic disease and comorbidities much more than non-allosteric proteins, and there is an association between being allosteric and involvement in disease; 2) they are enriched in many major disease types like hematopoietic diseases, cardiovascular diseases, cancers, diabetes, or diseases of the central nervous system; 3) variants from cancer genome-wide association studies are enriched near allosteric proteins, indicating their importance to polygenic traits; and 4) the importance of allosteric proteins in disease is due, at least partly, to their central positions in protein-protein interaction networks, and less due to their dynamical properties. György Abrusán, David B. Ascher, Michael Inouye |
PLoS Comput. Biol. | 3 |
| 2021 | Ten simple rules to make your computing more environmentally sustainableabstractRule 1: Calculate the carbon footprint of your workWe live in a world ruled by data, where a problem doesn't exist until it has been measured.There is still very limited information available about the carbon footprint of computational Loïc Lannelongue, Jason Grealey, Alex Bateman, Michael Inouye |
PLoS Comput. Biol. | 4 |
| 2021 | RANC: Reconfigurable Architecture for Neuromorphic ComputingabstractNeuromorphic architectures have been introduced as platforms for energy-efficient spiking neural network execution. The massive parallelism offered by these architectures has also triggered interest from nonmachine learning application domains. In order to lift the barriers to entry for hardware designers and application developers, we present RANC: a reconfigurable architecture for neuromorphic computing, an opensource highly flexible ecosystem that enables rapid experimentation with neuromorphic architectures in both software via C++ simulation and hardware via FPGA emulation. We present the utility of the RANC ecosystem by showing its ability to recreate behavior of IBM’s TrueNorth and validate with a direct comparison to IBM’s Compass simulation environment and published literature. RANC allows optimizing architectures based on application insights as well as prototyping future neuromorphic architectures that can support new classes of applications entirely. We demonstrate the highly parameterized and configurable nature of RANC by studying the impact of architectural changes on improving application mapping efficiency with quantitative analysis based on Alveo U250 FPGA. We present post routing resource usage and throughput analysis across implementations of synthetic aperture radar classification and vector matrix multiplication applications, and demonstrate a neuromorphic architecture that scales to emulating 259K distinct neurons and 73.3M distinct synapses. Joshua Mack, Ruben Purdy, Kris Rockowitz, Michael Inouye, Edward Richter, Spencer Valancius, Nirmal Kumbhare, Md Sahil Hassan, Kaitlin Lindsay Fair, John Mixter, Ali Akoglu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2019 | FastSpar: rapid and scalable correlation estimation for compositional dataabstractSUMMARY: A common goal of microbiome studies is the elucidation of community composition and member interactions using counts of taxonomic units extracted from sequence data. Inference of interaction networks from sparse and compositional data requires specialized statistical approaches. A popular solution is SparCC, however its performance limits the calculation of interaction networks for very high-dimensional datasets. Here we introduce FastSpar, an efficient and parallelizable implementation of the SparCC algorithm which rapidly infers correlation networks and calculates P-values using an unbiased estimator. We further demonstrate that FastSpar reduces network inference wall time by 2-3 orders of magnitude compared to SparCC. AVAILABILITY AND IMPLEMENTATION: FastSpar source code, precompiled binaries and platform packages are freely available on GitHub: github.com/scwatts/FastSpar. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Stephen C. Watts, Scott C. Ritchie, Michael Inouye, Kathryn E. Holt |
Bioinform. | 3 |
| 2017 | FlashPCA2: principal component analysis of Biobank-scale genotype datasetsabstractMOTIVATION: Principal component analysis (PCA) is a crucial step in quality control of genomic data and a common approach for understanding population genetic structure. With the advent of large genotyping studies involving hundreds of thousands of individuals, standard approaches are no longer feasible. However, when the full decomposition is not required, substantial computational savings can be made. RESULTS: We present FlashPCA2, a tool that can perform partial PCA on 1 million individuals faster than competing approaches, while requiring substantially less memory. AVAILABILITY AND IMPLEMENTATION: https://github.com/gabraham/flashpca . CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Gad Abraham, Michael Inouye |
Bioinform. | 3 |
| 2012 | SparSNP: Fast and memory-efficient analysis of all SNPs for phenotype predictionabstractBACKGROUND: A central goal of genomics is to predict phenotypic variation from genetic variation. Fitting predictive models to genome-wide and whole genome single nucleotide polymorphism (SNP) profiles allows us to estimate the predictive power of the SNPs and potentially develop diagnostic models for disease. However, many current datasets cannot be analysed with standard tools due to their large size. RESULTS: We introduce SparSNP, a tool for fitting lasso linear models for massive SNP datasets quickly and with very low memory requirements. In analysis on a large celiac disease case/control dataset, we show that SparSNP runs substantially faster than four other state-of-the-art tools for fitting large scale penalised models. SparSNP was one of only two tools that could successfully fit models to the entire celiac disease dataset, and it did so with superior performance. Compared with the other tools, the models generated by SparSNP had better than or equal to predictive performance in cross-validation. CONCLUSIONS: Genomic datasets are rapidly increasing in size, rendering existing approaches to model fitting impractical due to their prohibitive time or memory requirements. This study shows that SparSNP is an essential addition to the genomic analysis toolkit.SparSNP is available at http://www.genomics.csse.unimelb.edu.au/SparSNP. Gad Abraham, Adam Kowalczyk, Justin Zobel, Michael Inouye |
BMC Bioinform. | 4 |
| 2011 | Replication of epistatic DNA loci in two case-control GWAS studies using OPE algorithmabstractBackground One of the limiting factors of current genome-wide association studies (GWAS) is the inability of current methods to comprehensively examine SNP interactions for a reasonable sized dataset. It is hypothesised that this limitation is one of the reasons that GWAS studies have not been able to have a greater impact [1,2]. Many current methods for handling interactions are computationally expensive and do not scale to entire studies. Those methods that do scale often achieve this by pruning their datasets in some manner. This is commonly done by considering only those SNPs that show strong marginal effects, despite the fact that a strongly interacting pair may consist of SNPs with low effects individually. Benjamin Goudey, David Rawlinson 0001, Armita Zarnegar, Eder Kikianty, John Markham, Geoff MacIntyre, Gad Abraham, Linda Stern, Michael Inouye, Izhak Haviv, Adam Kowalczyk |
BMC Bioinform. | 10 |
| 2007 | A genotype calling algorithm for the Illumina BeadArray platformabstractMOTIVATION: Large-scale genotyping relies on the use of unsupervised automated calling algorithms to assign genotypes to hybridization data. A number of such calling algorithms have been recently established for the Affymetrix GeneChip genotyping technology. Here, we present a fast and accurate genotype calling algorithm for the Illumina BeadArray genotyping platforms. As the technology moves towards assaying millions of genetic polymorphisms simultaneously, there is a need for an integrated and easy-to-use software for calling genotypes. RESULTS: We have introduced a model-based genotype calling algorithm which does not rely on having prior training data or require computationally intensive procedures. The algorithm can assign genotypes to hybridization data from thousands of individuals simultaneously and pools information across multiple individuals to improve the calling. The method can accommodate variations in hybridization intensities which result in dramatic shifts of the position of the genotype clouds by identifying the optimal coordinates to initialize the algorithm. By incorporating the process of perturbation analysis, we can obtain a quality metric measuring the stability of the assigned genotype calls. We show that this quality metric can be used to identify SNPs with low call rates and accuracy. AVAILABILITY: The C++ executable for the algorithm described here is available by request from the authors. Yik Y. Teo, Michael Inouye, Kerrin S. Small, Rhian Gwilliam, Panagiotis Deloukas, Dominic Kwiatkowski, Taane G. Clark |
Bioinform. | 2 |