Michael Inouye

dblp:22/1580 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
4since 2021 · last 2024
0000-0001-9413-6520ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
4 papers
Bioinformatics and computational biology · 82% Computational science and engineering · 18%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Emerging computing paradigms · 70% Reconfigurable computing and FPGAs · 30%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computational science and engineering
benchmark framework
0.812024
Pitfalls of machine learning models for protein-protein interaction networks · Bioinform. 2024
Bioinformatics and computational biology
functional genomics
0.812024
Pitfalls of machine learning models for protein-protein interaction networks · Bioinform. 2024
Bioinformatics and computational biology › protein analysis › protein-protein interaction › protein-protein interaction network analysis
protein-protein interaction network inference
0.812024
Pitfalls of machine learning models for protein-protein interaction networks · Bioinform. 2024
Bioinformatics and computational biology
protein-protein interaction prediction
0.812024
Pitfalls of machine learning models for protein-protein interaction networks · Bioinform. 2024
Reconfigurable computing and FPGAs
FPGA-based emulation
0.512021
RANC: Reconfigurable Architecture for Neuromorphic Computing · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2021
Emerging computing paradigms
neuromorphic computing
0.512021
RANC: Reconfigurable Architecture for Neuromorphic Computing · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2021
Emerging computing paradigms › neuromorphic computing
spiking neural network
0.512021
RANC: Reconfigurable Architecture for Neuromorphic Computing · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2021
Bioinformatics and computational biology › computational microbiology
microbiome analysis
0.412019
FastSpar: rapid and scalable correlation estimation for compositional data · Bioinform. 2019
Bioinformatics and computational biology › population genetics
ancestry inference
0.312017
FlashPCA2: principal component analysis of Biobank-scale genotype datasets · Bioinform. 2017
Bioinformatics and computational biology
population genetics
0.312017
FlashPCA2: principal component analysis of Biobank-scale genotype datasets · Bioinform. 2017
Emerging computing paradigms
neuromorphic hardware
0.112021
RANC: Reconfigurable Architecture for Neuromorphic Computing · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2021
Algorithms and data structures
parallel algorithms
0.112019
FastSpar: rapid and scalable correlation estimation for compositional data · Bioinform. 2019
Bioinformatics and computational biology › statistical genetics
genotype calling
0.112007
A genotype calling algorithm for the Illumina BeadArray platform · Bioinform. 2007
Bioinformatics and computational biology › genomics
genotyping
0.112007
A genotype calling algorithm for the Illumina BeadArray platform · Bioinform. 2007

Methods — techniques the papers use, named apart from their topics

unbiased estimator · 0.8reproducibility analysis · 0.8parallelization · 0.8machine learning · 0.8benchmarking · 0.8SparCC · 0.8c++ simulation · 0.5FPGA emulation · 0.5randomized algorithm · 0.3partial singular value decomposition · 0.3quality metric · 0.1perturbation analysis · 0.1
YearPublicationVenuePosition
2024 Pitfalls of machine learning models for protein-protein interaction networks
abstract
MOTIVATION: Protein-protein interactions (PPIs) are essential to understanding biological pathways as well as their roles in development and disease. Computational tools, based on classic machine learning, have been successful at predicting PPIs in silico, but the lack of consistent and reliable frameworks for this task has led to network models that are difficult to compare and discrepancies between algorithms that remain unexplained. RESULTS: To better understand the underlying inference mechanisms that underpin these models, we designed an open-source framework for benchmarking that accounts for a range of biological and statistical pitfalls while facilitating reproducibility. We use it to shed light on the impact of network topology and how different algorithms deal with highly connected proteins. By studying functional genomics-based and sequence-based models on human PPIs, we show their complementarity as the former performs best on lone proteins while the latter specializes in interactions involving hubs. We also show that algorithm design has little impact on performance with functional genomic data. We replicate our results between both human and S. cerevisiae data and demonstrate that models using functional genomics are better suited to PPI prediction across species. With rapidly increasing amounts of sequence and functional genomics data, our study provides a principled foundation for future construction, comparison, and application of PPI networks. AVAILABILITY AND IMPLEMENTATION: The code and data are available on GitHub: https://github.com/Llannelongue/B4PPI.
Loïc Lannelongue, Michael Inouye
Bioinform.2
2022 Known allosteric proteins have central roles in genetic disease
abstract
Allostery is a form of protein regulation, where ligands that bind sites located apart from the active site can modify the activity of the protein. The molecular mechanisms of allostery have been extensively studied, because allosteric sites are less conserved than active sites, and drugs targeting them are more specific than drugs binding the active sites. Here we quantify the importance of allostery in genetic disease. We show that 1) known allosteric proteins are central in disease networks, contribute to genetic disease and comorbidities much more than non-allosteric proteins, and there is an association between being allosteric and involvement in disease; 2) they are enriched in many major disease types like hematopoietic diseases, cardiovascular diseases, cancers, diabetes, or diseases of the central nervous system; 3) variants from cancer genome-wide association studies are enriched near allosteric proteins, indicating their importance to polygenic traits; and 4) the importance of allosteric proteins in disease is due, at least partly, to their central positions in protein-protein interaction networks, and less due to their dynamical properties.
György Abrusán, David B. Ascher, Michael Inouye
PLoS Comput. Biol.3
2021 Ten simple rules to make your computing more environmentally sustainable
abstract
Rule 1: Calculate the carbon footprint of your workWe live in a world ruled by data, where a problem doesn't exist until it has been measured.There is still very limited information available about the carbon footprint of computational
Loïc Lannelongue, Jason Grealey, Alex Bateman, Michael Inouye
PLoS Comput. Biol.4
2021 RANC: Reconfigurable Architecture for Neuromorphic Computing
abstract
Neuromorphic architectures have been introduced as platforms for energy-efficient spiking neural network execution. The massive parallelism offered by these architectures has also triggered interest from nonmachine learning application domains. In order to lift the barriers to entry for hardware designers and application developers, we present RANC: a reconfigurable architecture for neuromorphic computing, an opensource highly flexible ecosystem that enables rapid experimentation with neuromorphic architectures in both software via C++ simulation and hardware via FPGA emulation. We present the utility of the RANC ecosystem by showing its ability to recreate behavior of IBM’s TrueNorth and validate with a direct comparison to IBM’s Compass simulation environment and published literature. RANC allows optimizing architectures based on application insights as well as prototyping future neuromorphic architectures that can support new classes of applications entirely. We demonstrate the highly parameterized and configurable nature of RANC by studying the impact of architectural changes on improving application mapping efficiency with quantitative analysis based on Alveo U250 FPGA. We present post routing resource usage and throughput analysis across implementations of synthetic aperture radar classification and vector matrix multiplication applications, and demonstrate a neuromorphic architecture that scales to emulating 259K distinct neurons and 73.3M distinct synapses.
Joshua Mack, Ruben Purdy, Kris Rockowitz, Michael Inouye, Edward Richter, Spencer Valancius, Nirmal Kumbhare, Md Sahil Hassan, Kaitlin Lindsay Fair, John Mixter, Ali Akoglu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2019 FastSpar: rapid and scalable correlation estimation for compositional data
abstract
SUMMARY: A common goal of microbiome studies is the elucidation of community composition and member interactions using counts of taxonomic units extracted from sequence data. Inference of interaction networks from sparse and compositional data requires specialized statistical approaches. A popular solution is SparCC, however its performance limits the calculation of interaction networks for very high-dimensional datasets. Here we introduce FastSpar, an efficient and parallelizable implementation of the SparCC algorithm which rapidly infers correlation networks and calculates P-values using an unbiased estimator. We further demonstrate that FastSpar reduces network inference wall time by 2-3 orders of magnitude compared to SparCC. AVAILABILITY AND IMPLEMENTATION: FastSpar source code, precompiled binaries and platform packages are freely available on GitHub: github.com/scwatts/FastSpar. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Stephen C. Watts, Scott C. Ritchie, Michael Inouye, Kathryn E. Holt
Bioinform.3
2017 FlashPCA2: principal component analysis of Biobank-scale genotype datasets
abstract
MOTIVATION: Principal component analysis (PCA) is a crucial step in quality control of genomic data and a common approach for understanding population genetic structure. With the advent of large genotyping studies involving hundreds of thousands of individuals, standard approaches are no longer feasible. However, when the full decomposition is not required, substantial computational savings can be made. RESULTS: We present FlashPCA2, a tool that can perform partial PCA on 1 million individuals faster than competing approaches, while requiring substantially less memory. AVAILABILITY AND IMPLEMENTATION: https://github.com/gabraham/flashpca . CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Gad Abraham, Michael Inouye
Bioinform.3
2012 SparSNP: Fast and memory-efficient analysis of all SNPs for phenotype prediction
abstract
BACKGROUND: A central goal of genomics is to predict phenotypic variation from genetic variation. Fitting predictive models to genome-wide and whole genome single nucleotide polymorphism (SNP) profiles allows us to estimate the predictive power of the SNPs and potentially develop diagnostic models for disease. However, many current datasets cannot be analysed with standard tools due to their large size. RESULTS: We introduce SparSNP, a tool for fitting lasso linear models for massive SNP datasets quickly and with very low memory requirements. In analysis on a large celiac disease case/control dataset, we show that SparSNP runs substantially faster than four other state-of-the-art tools for fitting large scale penalised models. SparSNP was one of only two tools that could successfully fit models to the entire celiac disease dataset, and it did so with superior performance. Compared with the other tools, the models generated by SparSNP had better than or equal to predictive performance in cross-validation. CONCLUSIONS: Genomic datasets are rapidly increasing in size, rendering existing approaches to model fitting impractical due to their prohibitive time or memory requirements. This study shows that SparSNP is an essential addition to the genomic analysis toolkit.SparSNP is available at http://www.genomics.csse.unimelb.edu.au/SparSNP.
Gad Abraham, Adam Kowalczyk, Justin Zobel, Michael Inouye
BMC Bioinform.4
2011 Replication of epistatic DNA loci in two case-control GWAS studies using OPE algorithm
abstract
Background One of the limiting factors of current genome-wide association studies (GWAS) is the inability of current methods to comprehensively examine SNP interactions for a reasonable sized dataset. It is hypothesised that this limitation is one of the reasons that GWAS studies have not been able to have a greater impact [1,2]. Many current methods for handling interactions are computationally expensive and do not scale to entire studies. Those methods that do scale often achieve this by pruning their datasets in some manner. This is commonly done by considering only those SNPs that show strong marginal effects, despite the fact that a strongly interacting pair may consist of SNPs with low effects individually.
Benjamin Goudey, David Rawlinson 0001, Armita Zarnegar, Eder Kikianty, John Markham, Geoff MacIntyre, Gad Abraham, Linda Stern, Michael Inouye, Izhak Haviv, Adam Kowalczyk
BMC Bioinform.10
2007 A genotype calling algorithm for the Illumina BeadArray platform
abstract
MOTIVATION: Large-scale genotyping relies on the use of unsupervised automated calling algorithms to assign genotypes to hybridization data. A number of such calling algorithms have been recently established for the Affymetrix GeneChip genotyping technology. Here, we present a fast and accurate genotype calling algorithm for the Illumina BeadArray genotyping platforms. As the technology moves towards assaying millions of genetic polymorphisms simultaneously, there is a need for an integrated and easy-to-use software for calling genotypes. RESULTS: We have introduced a model-based genotype calling algorithm which does not rely on having prior training data or require computationally intensive procedures. The algorithm can assign genotypes to hybridization data from thousands of individuals simultaneously and pools information across multiple individuals to improve the calling. The method can accommodate variations in hybridization intensities which result in dramatic shifts of the position of the genotype clouds by identifying the optimal coordinates to initialize the algorithm. By incorporating the process of perturbation analysis, we can obtain a quality metric measuring the stability of the assigned genotype calls. We show that this quality metric can be used to identify SNPs with low call rates and accuracy. AVAILABILITY: The C++ executable for the algorithm described here is available by request from the authors.
Yik Y. Teo, Michael Inouye, Kerrin S. Small, Rhian Gwilliam, Panagiotis Deloukas, Dominic Kwiatkowski, Taane G. Clark
Bioinform.2