VLDB 2026 Research / reviewers in the wild / expert
Scott J. Emrich
dblp:76/2937
· DBLP profile ↗
33ranked-venue papers
1as first author
4since 2021 · last 2024
0000-0002-5741-4517ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 21 · 1 first-author · 4 since 2021Systems, architecture and hardware · 12Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | CodonT5: A Multi-task Codon Language Model to Perform Generative Codon OptimizationabstractCodon language modeling is an emerging area of research interest. Prior efforts in codon optimization have mostly focused on the effect codon preferences have on protein translation speed either at single codons (e.g., selection preference measures) or entire gene sequences (e.g., Codon Adaptation Index and related estimates of gene expression). Model parameters are almost always trained based on a specific organism/task. Here we explore a first ever look at the impact of multi-task model pre-training with the end goal of performing codon "translations" between species. Significantly, our model, CodonT5, is able to accurately predict across multiple tasks and generate comparable codon optimizations to previously established methods using only the reference sequences without being necessarily dependent on any pre-computed metrics of codon preference. This is important as it paves the way for more diverse sequence-to-sequence modeling that will be necessary in many applications involving replicating a reference protein in a host (e.g., mRNA vaccines). Ashley Babjac, Scott J. Emrich |
BIBM | 2 |
| 2023 | Adapting Protein Language Models for Explainable Fine-Grained Evolutionary Pattern DiscoveryabstractOrganisms that live in different environments face different evolutionary pressures. As such, organisms that have more successful phenotypes reproduce more frequently, but differing selective pressures acting at the organismal level can influence genes, and thus proteins. Understanding how proteins adapt across environments may therefore be useful in engineering proteins for specific environments as well as to improve our understanding of basic biology. In this work, we explicitly compare homologous (read: paired) proteins from different environments. While previous studies have explored the relevant evolutionary pressures in one of these environments [11], [17] and genomic responses to those pressures [1], [28], no prior computational study of their proteins has been performed. We apply ESM-2 [20] and although there is no signal in our negative control (two divergent yeast strains) as expected, we obtain near perfect prediction accuracy for our selected environmental gradient–the well-established subsurface vs. surface biome. We further show that ESM-2 is able to capture relevant fine-grained biological patterns in its embedding space, even in its smallest model. Significantly, we demonstrate that these embeddings can be interpreted using a novel visualization pipeline built using explainable AI techniques. Ashley Babjac, Owen Queen, Shawn-Patrick Barhorst, Kambiz Kalhor, Andrew D. Steen, Scott J. Emrich |
BIBM | 6 |
| 2021 | Fine-Grained Synonymous Codon Usage Patterns and their Potential Role in Functional Protein ProductionabstractMultiple computational models exist that leverage the idea that some codons are “preferred” to others when producing functional proteins. In this work, we aim to address the following two questions: (i) do proteins have similar codon usage subprofiles?, (ii) if so, can they be associated to specific Gene Ontology (GO) terms? To answer these questions, we analyze two different data sets E. coli and S. cerevisiae (yeast). For both, we use codon usage profiles that have been linked to protein co-translational folding and perform unsupervised clustering. Significantly, our clusters determine biologically meaningful patterns consistent with previously published observations. Our results suggest that codon patterns have potential to be incorporated into more complex deep learning-inspired frameworks in future work to better understand and predict functional protein production in a cell. Ashley Babjac, Jun Li 0052, Scott J. Emrich |
BIBM | 3 |
| 2021 | LASSO-based feature selection for improved microbial and microbiome classificationabstractMicrobial and microbiome data consist of large data matrices that contain many irrelevant or colinear features. This can hinder the performance of machine learning algorithms because of the “curse of dimensionality.” Here, we attempted to overcome this potential limitation by employing feature selection, specifically a custom LASSO-based framework. We tested our approach by performing classification tasks on two datasets, one consisting of microbiome diversity data from poplar trees and one of gene presence/absence data from E. coli samples taken from hospitalized patients at risk of sepsis. Through a comparison to existing feature selection methods, we found that LASSO consistently outperformed on several key classification metrics, most notably AUC. On several prediction tasks, LASSO feature selection was able to increase AUC over the baseline method from around 0.5 to near-perfect values (0.99 and above). Our results suggest that our new LASSO framework can generate more meaningful features for classification algorithms relative to similar feature selection methods. Owen Queen, Scott J. Emrich |
BIBM | 2 |
| 2020 | Extreme Phenotype Sampling Improves LASSO and Random Forest Marker Selection for Complex TraitsabstractMost attempts to fit a supervised machine learning (ML) model in bioinformatics try to predict the full range of trait or response values. While such prediction tasks effectively capture the entire phenotypic range of the samples, they are cost prohibitive and can be statistically underpowered for detection of rare variants. In a study design known as extreme phenotype sampling (EPS), samples are selected from the two extremes of the phenotypic distribution. This approach is costcutting, by reducing genotyping/sequencing costs, as well as capable of increasing statistical power. Although combining EPS with ML algorithms has the potential to enhance association studies by improving their computational efficiency, EPS-ML approaches have seen limited use. In this paper we demonstrate an efficient and effective approach to leverage the EPS study design using LASSO regression and random forests, two commonly used ML algorithms within the broader bioinformatics community. We analyze two distinct data sets: leaf expression values generated from black cottonwood and malaria parasite transcriptome data collected from patients. We demonstrate that focusing only on the phenotypic extremes of these sample sets (by forming binary classes) can select more biologically meaningful features than using the full range. This approach will be useful to investigators when examining complex or novel traits. It is particularly well-suited to RNA-seq data where investigators often want to narrow attention to a small number of candidate transcripts out of a large initial pool. Our approach intentionally leverages existing software with efficient implementations to enable future applications of EPS-ML. Cai W. John, Wellington Muchero, Scott J. Emrich |
BIBM | 3 |
| 2020 | Network analysis of synonymous codon usageabstractMOTIVATION: Most amino acids are encoded by multiple synonymous codons, some of which are used more rarely than others. Analyses of positions of such rare codons in protein sequences revealed that rare codons can impact co-translational protein folding and that positions of some rare codons are evolutionarily conserved. Analyses of their positions in protein 3-dimensional structures, which are richer in biochemical information than sequences alone, might further explain the role of rare codons in protein folding. RESULTS: We model protein structures as networks and use network centrality to measure the structural position of an amino acid. We first validate that amino acids buried within the structural core are network-central, and those on the surface are not. Then, we study potential differences between network centralities and thus structural positions of amino acids encoded by conserved rare, non-conserved rare and commonly used codons. We find that in 84% of proteins, the three codon categories occupy significantly different structural positions. We examine protein groups showing different codon centrality trends, i.e. different relationships between structural positions of the three codon categories. We see several cases of all proteins from our data with some structural or functional property being in the same group. Also, we see a case of all proteins in some group having the same property. Our work shows that codon usage is linked to the final protein structure and thus possibly to co-translational protein folding. AVAILABILITY AND IMPLEMENTATION: https://nd.edu/∼cone/CodonUsage/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Khalique Newaz, Gabriel Wright, Jacob Piland, Jun Li 0052, Patricia L. Clark, Scott J. Emrich, Tijana Milenkovic |
Bioinform. | 6 |
| 2019 | Highly Accurate and Efficient Data-Driven Methods for Genotype ImputationabstractHigh-throughput sequencing techniques have generated massive quantities of genotype data. Haplotype phasing has proven to be a useful and effective method for analyzing these data. However, the quality of phasing is undermined due to missing information. Imputation provides an effective means of improving the underlying genotype information. For model organisms, imputation can rely on an available reference genotype panel and a physical or genetic map. For non-model organisms, which often do not have a genotype panel, it is important to design an imputation technique that does not rely on reference data. Here, we present Accurate Data-Driven Imputation Technique (ADDIT), which is composed of two data-driven algorithms capable of handling data generated from model and non-model organisms. The non-model variant of ADDIT (referred to as ADDIT-NM) employs statistical inference methods to impute missing genotypes, whereas the model variant (referred to as ADDIT-M) leverages a supervised learning-based approach for imputation. We demonstrate that both variants of ADDIT are more accurate, faster, and require less memory than leading state-of-the-art imputation tools using model (human) and non-model (maize, apple, and grape) genotype data. Software Availability: The source code of ADDIT and test data sets are available at https://github.com/NDBL/ADDIT. Olivia Choudhury, Ankush Chakrabarty, Scott J. Emrich |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2018 | Effects from structure of Metabarcode Sequences on Lossy Analysis of Microbiome Data
David C. Molik, Michael E. Pfrender, Scott J. Emrich |
BIBM | 3 |
| 2018 | The Effects of Normalization, Transformation, and Rarefaction on Clustering of OTU Abundance
DeAndre A. Tomlinson, David C. Molik, Michael E. Pfrender, Scott J. Emrich |
BIBM | 4 |
| 2018 | Predicting Local Inversions Using Rectangle Clustering and Representative Rectangle Prediction
Shenglong Zhu, Scott J. Emrich, Danny Ziyi Chen |
BIBM | 2 |
| 2018 | Combining Static and Dynamic Storage Management for Data Intensive Scientific WorkflowsabstractWorkflow management systems are widely used to express and execute highly parallel applications. For data-intensive workflows, storage can be the constraining resource: The number of tasks running at once must be artificially limited to not overflow the space available in the filesystem. It is all too easy for a user to dispatch a workflow which consumes all available storage and disrupts all system users. To address these issues, we present a three-tiered approach to workflow storage management: (1) A static analysis algorithm which analyzes the storage needs of a workflow before execution, giving a realistic prediction of success or failure. (2) An online storage management algorithm which accounts for the storage needed by future tasks to avoid deadlock at runtime. (3) A task containment system which limits storage consumption of individual tasks, enabling the strong guarantees of the static analysis and dynamic management algorithms. We demonstrate the application of these techniques on three complex workflows. Nicholas L. Hazekamp, Nathaniel Kremer-Herman, Benjamín Tovar, Haiyan Meng, Olivia Choudhury, Scott J. Emrich, Douglas Thain |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2017 | Stable feature ranking with logistic regression ensemblesabstractBeyond automated classification, supervised machine-learning models can be interpreted to find which features or combination of features distinguish sets of classes. Logistic Regression (LR) is an example of a model well-suited for human interpretation. Unfortunately, results from feature ranking with LR may not be reliable and reproducible for the same dataset. We demonstrate that stability and consistency can be achieved via ensembles (“LR ensembles”). As a specific example of the real-world utility of our associated framework, we apply LR ensembles to single-nucleotide polymorphisms (SNPs) associated with the recent speciation of the malaria vectors Anopheles gambiae and Anopheles coluzzii and compare with the more common univariate metric FST. Ronald J. Nowling, Scott J. Emrich |
BIBM | 2 |
| 2017 | Inversion detection using PacBio long readsabstractStructural variation is important in disease etiology and ecological adaptation. Prior work has focused on using either only short paired-end reads or a hybrid approach that combines long and short reads to detect structural variants. Few methods have focused solely on using long reads. Here, we aim to detect a specific type of structural variation, large inversions, using only raw PacBio long reads. We propose a new breakpoint detection approach that is complementary to current state-of-the-art methods and models inversion detection as a Max-Cut problem. We show that this new approach is powerful for detecting large inversions when compared with popular structural variation detection tools. In particular, it yields high Positive Predictive Value (PPV), relatively good sensitivity for detecting large inversions, and shows potential for detecting inversions in more complex genomes. Our software is publicly available at https://bitbucket.org/NDBL/invdet. Shenglong Zhu, Scott J. Emrich, Danny Ziyi Chen |
BIBM | 2 |
| 2017 | Widespread position-specific conservation of synonymous rare codons within coding sequencesabstractSynonymous rare codons are considered to be sub-optimal for gene expression because they are translated more slowly than common codons. Yet surprisingly, many protein coding sequences include large clusters of synonymous rare codons. Rare codons at the 5' terminus of coding sequences have been shown to increase translational efficiency. Although a general functional role for synonymous rare codons farther within coding sequences has not yet been established, several recent reports have identified rare-to-common synonymous codon substitutions that impair folding of the encoded protein. Here we test the hypothesis that although the usage frequencies of synonymous codons change from organism to organism, codon rarity will be conserved at specific positions in a set of homologous coding sequences, for example to tune translation rate without altering a protein sequence. Such conservation of rarity-rather than specific codon identity-could coordinate co-translational folding of the encoded protein. We demonstrate that many rare codon cluster positions are indeed conserved within homologous coding sequences across diverse eukaryotic, bacterial, and archaeal species, suggesting they result from positive selection and have a functional role. Most conserved rare codon clusters occur within rather than between conserved protein domains, challenging the view that their primary function is to facilitate co-translational folding after synthesis of an autonomous structural unit. Instead, many conserved rare codon clusters separate smaller protein structural motifs within structural domains. These smaller motifs typically fold faster than an entire domain, on a time scale more consistent with translation rate modulation by synonymous codon usage. While proteins with conserved rare codon clusters are structurally and functionally diverse, they are enriched in functions associated with organism growth and development, suggesting an important role for synonymous codon usage in organism physiology. The identification of conserved rare codon clusters advances our understanding of distinct, functional roles for otherwise synonymous codons and enables experimental testing of the impact of synonymous codon usage on the production of functional proteins. Julie L. Chaney, Aaron Steele, Rory Carmichael, Anabel Rodriguez, Alicia T. Specht, Kim Ngo, Jun Li 0052, Scott J. Emrich, Patricia L. Clark |
PLoS Comput. Biol. | 8 |
| 2015 | A computational framework for integrative analysis of large microbial genomics dataabstractThe availability of huge amount of genome sequence data from natural microbial consortia enables integrated analysis to resolve the genetic and metabolic potential of microbial communities, to establish how functions are partitioned in and among populations, and to reveal how microbial communities evolve and adapt across multiple environments. In this paper, we propose to analyze comparative microbial genomes using a computational framework. The framework is designed to investigate genome context patterns of microbial diversity. With an application to investigate functions of three environments (human gut, soil, and marine), we demonstrated that the developed computational framework was able to identify functional modules and evaluate the functional roles of those modules in microbial communities as response to environmental change. We found different gene networks among microbial communities living in different environments. We showed that modules identified by our framework can be computationally annotated to study their biological functions. Erliang Zeng, Wei Zhang 0175, Scott J. Emrich, Joshua Livermore, Stuart E. Jones |
BIBM | 3 |
| 2015 | Balancing Thread-Level and Task-Level Parallelism for Data-Intensive Workloads on Clusters and CloudsabstractThe runtime configuration of parallel and distributed applications remains a mysterious art. To tune an application on a particular system, the end-user must choose the number of machines, the number of cores per task, the data partitioning strategy, and so on, all of which result in a combinatorial explosion of choices. While one might try to exhaustively evaluate all choices in search of the optimal, the end user's goal is simply to run the application once with reasonable performance by avoiding terrible configurations. To address this problem, we present a hybrid technique based on regression models for tuning data intensive bioinformatics applications: the sequential computational kernel is characterized empirically and then incorporated into an ab initio model of the distributed system. We demonstrate this technique on the commonly-used applications BWA, Bowtie2, and BLASR and validate the accuracy of our proposed models on clouds and clusters. Olivia Choudhury, Dinesh Rajan, Nicholas L. Hazekamp, Sandra Gesing, Douglas Thain, Scott J. Emrich |
CLUSTER | 6 |
| 2015 | Scaling Up Bioinformatics Workflows with Dynamic Job Expansion: A Case Study Using Galaxy and MakeflowabstractLogical workflow management systems provide a user-friendly portal through which data can be processed using a sequence of standard tools. These logical workflows are a natural way to express the high level intent of the user, and to share the structure and the results with other users. However, logical workflows are not necessarily suited to expressing parallelism for very large runs. As the amount of data is scaled up, the run time of each node in the logical workflow may become extreme. We propose a technique of job expansion to solve this problem. When job expansion is applied to a logical workflow, each node in the workflow is itself expanded into a large performance workflow that may consist of hundreds to thousands of tasks that can be executed in parallel, thus enabling high concurrency and scalability. From the user's perspective, nothing has changed and the logical workflow remains in its original form. To demonstrate this technique, we have applied job expansion to a selection of bioinformatics applications running in the Galaxy workflow management system. Each job in the workflow is expanded into a highly parallel workflow executed using Makeflow, which is well suited to express high levels of parallelism. Work Queue is then utilized for execution because of its ability to quickly dispatch tasks and cache files for later reuse. After applying job expansion, we improve the execution time of BWA 18X and GATK 402X, with a total speedup of 61.5X on the workflow. We also take a look at the systems behavior since its launch to analyze its effectiveness. Nicholas L. Hazekamp, Joseph Sarro, Olivia Choudhury, Sandra Gesing, Scott J. Emrich, Douglas Thain |
e-Science | 5 |
| 2015 | RNA-Rocket: an RNA-Seq analysis resource for infectious disease researchabstractMOTIVATION: RNA-Seq is a method for profiling transcription using high-throughput sequencing and is an important component of many research projects that wish to study transcript isoforms, condition specific expression and transcriptional structure. The methods, tools and technologies used to perform RNA-Seq analysis continue to change, creating a bioinformatics challenge for researchers who wish to exploit these data. Resources that bring together genomic data, analysis tools, educational material and computational infrastructure can minimize the overhead required of life science researchers. RESULTS: RNA-Rocket is a free service that provides access to RNA-Seq and ChIP-Seq analysis tools for studying infectious diseases. The site makes available thousands of pre-indexed genomes, their annotations and the ability to stream results to the bioinformatics resources VectorBase, EuPathDB and PATRIC. The site also provides a combination of experimental data and metadata, examples of pre-computed analysis, step-by-step guides and a user interface designed to enable both novice and experienced users of RNA-Seq data. AVAILABILITY AND IMPLEMENTATION: RNA-Rocket is available at rnaseq.pathogenportal.org. Source code for this project can be found at github.com/cidvbi/PathogenPortal. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary materials are available at Bioinformatics online. Andrew S. Warren, Cristina Aurrecoechea, Brian P. Brunk, Prerak Desai, Scott J. Emrich, Gloria I. Giraldo-Calderón, Omar S. Harb, Deborah Hix, Daniel Lawson, Dustin Machi, Chunhong Mao, Michael McClelland, Eric K. Nordberg, Maulik Shukla, Leslie B. Vosshall, Alice R. Wattam, Rebecca Will, Hyun Seung Yoo, Bruno W. S. Sobral |
Bioinform. | 5 |
| 2014 | Accelerating Comparative Genomics Workflows in a Distributed Environment with Optimized Data PartitioningabstractThe advent of new sequencing technology has generated massive amounts of biological data at unprecedented rates. High-throughput bioinformatics tools are required to keep pace with this. Here, we implement a workflow-based model for parallelizing the data intensive task of genome alignment and variant calling with BWA and GATK's Haplotype Caller. We explore different approaches of partitioning data and how each affect the run time. We observe granularity-based partitioning for BWA and alignment-based partitioning for Halo type Caller to be the optimal choices for the pipeline. We identify the various challenges encountered while developing such an application and provide an insight into addressing them. We report significant performance improvements, from 12 days to 4 hours, while running the BWA-GATK pipeline using 100 nodes for analyzing high-coverage oak tree data. Olivia Choudhury, Nicholas L. Hazekamp, Douglas Thain, Scott J. Emrich |
CCGRID | 4 |
| 2014 | Expanding Tasks of Logical Workflows Into Independent Workflows for Improved ScalabilityabstractWorkflow Management Systems, such as Galaxy and Taverna, provide a portal through which data can be processed using a sequence of different tools. This sequence allows for the creation of a logical workflow that describes the process. However, when the data workload becomes large enough the time spent in each logical step increases making it difficult to run the workflow fast and efficiently. The proposed solutions is to use task level expansion. Task expansion aims to take each step of the logical workflow and expand it into a new self-contained workflow. These workflows would allow for greater scalability and concurrency by creating more tasks. The resulting workflows will be used indistinguishably from the original tool, but perform more quickly and efficiently. The concept was applied to the BWA tool in Galaxy and we were able to see a 7.36 times speedup in runtime on our 32 GB dataset. Nicholas L. Hazekamp, Olivia Choudhury, Sandra Gesing, Scott J. Emrich, Douglas Thain |
CCGRID | 4 |
| 2014 | Adapting bioinformatics applications for heterogeneous systems: a case studyabstractSUMMARY The advent of new sequencing technologies has generated extremely large amounts of information. To successfully apply bioinformatics tools to such large datasets, they need to exhibit scalability and ideally elasticity in diverse computing environments. We describe the application of previously obtained lessons to a new workflow with and without shared file storage. Because the original workflows have an intractable sequential running times on large datasets, we propose lessons and results for refactoring bioinformatics tools for elastic scaling on personal clouds. Our case studies describe the various challenges faced when constructing such a workflow, from dealing with failure detection, to managing dependencies, to handling the quirks of the underlying operating systems. The practice of scaling bioinformatics tools is increasingly commonplace. As such, this hands‐on application of refactoring techniques can serve as a valuable guide. Significantly, our customized Makeflow framework enabled generalizable deployment on a wider variety of systems while substantially reducing wall clock runtimes using hundreds of cores. Copyright © 2012 John Wiley & Sons, Ltd. Irena Lanc, Peter Bui, Douglas Thain, Scott J. Emrich |
Concurr. Comput. Pract. Exp. | 4 |
| 2013 | Case Studies in Designing Elastic ApplicationsabstractClusters, clouds, and grids offer access to large scale computational resources at low cost. This is especially appealing to scientific applications that require a very large scale to compete in the research space. However, the resources available across these platforms differ significantly in their availability, hardware, environment, performance, cost of use, and more. This requires the use of elastic applications that can adapt to the resources available at run-time, transparently handling heterogeneity and failures. In this paper, we present case studies of several elastic applications built using the Work Queue programming framework. From this experience, we offer six general guidelines for the design and implementation of elastic applications that run on thousands of processors. Dinesh Rajan, Andrew Thrasher, Badi Abdul-Wahid, Jesús A. Izaguirre, Scott J. Emrich, Douglas Thain |
CCGRID | 5 |
| 2012 | A Framework for Scalable Genome Assembly on Clusters, Clouds, and GridsabstractBioinformatics researchers need efficient means to process large collections of genomic sequence data. One application of interest, genome assembly, has great potential for parallelization; however, most previous attempts at parallelization require uncommon high-end hardware. This paper introduces the Scalable Assembler at Notre Dame (SAND) framework that can achieve significant speedup using large numbers of commodity machines harnessed from clusters, clouds, and grids. SAND interfaces with the Celera open-source assembly toolkit, replacing two independent sequential modules with scalable parallel alternatives: the candidate selector exploits distributed memory capacity, and the sequence aligner exploits distributed computing capacity. For large problems, these modules provide robust task and data management while also achieving speedup with high efficiency. We show results for several data sets ranging from 738 thousand to over 320 million alignments using resources ranging from a small cluster to more than a thousand nodes spanning three institutions. Christopher Moretti, Andrew Thrasher, Michael Olson, Scott J. Emrich, Douglas Thain |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2011 | Biocompute 2.0: an improved collaborative workspace for data intensive bio-scienceabstractSUMMARY The explosion of data in the biological community requires scalable and flexible portals for bioinformatics. To help address this need, we proposed characteristics needed for rigorous, reproducible, and collaborative resources for data‐intensive science. Implementing a system with these characteristics exposed challenges in user interface, data distribution, and workflow description/execution. We describe ongoing responses to these and other challenges. Our Data‐Action‐Queue design pattern addresses user interface and system organization concepts. A dynamic data distribution mechanism lays the foundation for the management of persistent datasets. Makeflow facilitates the simple description and execution of complex multi‐part jobs and forms the kernel of a module system powering diverse bioinformatics applications. Our improved web portal, Biocompute 2.0, has been in production use since the summer of 2010. Through it and its predecessor, we have provided over 56 years of CPU time through its five modules—BLAST, SSAHA, SHRIMP, BWA, and SNPEXP—to research groups at three universities. In this paper, we describe the goals and interface to the system, its architecture and performance, and the insights gained in its development. Copyright © 2011 John Wiley & Sons, Ltd. Rory Carmichael, Patrick Braga-Henebry, Douglas Thain, Scott J. Emrich |
Concurr. Comput. Pract. Exp. | 4 |
| 2010 | A supervised learning approach to the unsupervised clustering of genesabstractClustering is a common step in the analysis of microarray data. Microarrays enable simultaneous high-throughput measurement of the expression level of genes. These data can be used to explore relationships between genes and can guide development of drugs and further research. A typical first step in the analysis of these data is to use an agglomerative hierarchical clustering algorithm on the correlation between all gene pairs. While this simple approach has been successful it fails to identify many genetic interactions that may be important for drug design and other important applications. We present an approach to the clustering of expression data that utilizes known gene-gene interaction data to improve results for already commonly used clustering techniques. The approach creates an ensemble similarity measure that can be used as input to common clustering techniques and provides results with increased biological significance while not altering the clustering approach at all. Andrew K. Rider, Geoffrey Siwo, Scott J. Emrich, Michael T. Ferdig, Nitesh V. Chawla |
BIBM | 3 |
| 2010 | A two-stage machine learning approach for pathway analysisabstractAnalysis of gene expression data has emerged as an important approach to discover active pathways related to biological phenotypes. Previous pathway analysis methods use all genes in a pathway for linking it to a particular phenotype. Using only a subset of informative genes, however, could better classify samples. Here, we propose a two-stage machine learning approach for pathway analysis. During the first stage, informative genes that can represent a pathway are selected using feature selection methods. These “representative genes” are mostly associated with the phenotype of interest. In the second stage, pathways are ranked based on their “representative genes” using classification methods. We applied our two-stage approach on three gene expression datasets. The results indicate our method does outperform methods that consider every gene in a pathway. Wei Zhang 0175, Scott J. Emrich, Erliang Zeng |
BIBM | 2 |
| 2010 | Biocompute: towards a collaborative workspace for data intensive bio-scienceabstractThe explosion of data in the biological community demands the development of more scalable and flexible portals for bioinformatic computation. To address this need, we put forth characteristics needed for rigorous, reproducible, and collaborative resources for data intensive science. Implementing a system with these characteristics exposed challenges in user interface, data distribution, and workflow description/execution. We describe several responses to these challenges. The Data-Action-Queue metaphor addresses user interface and system organization concepts. A dynamic data distribution mechanism lays the foundation for the management of persistent datasets. The Makeflow workflow facilitates the simple description and execution of complex multipart jobs. The resulting web portal, Biocompute, has been in production use at the University of Notre Dame's Bioinformatics Core Facility since the summer of 2009. It has provided over seven years of CPU time through its three sequence search modules --- BLAST, SSAHA, and SHRIMP --- to ten biological and bioinformatic research groups spanning three universities. In this paper we describe the goals and interface to the system, its architecture and performance, and the insights gained in its development. Rory Carmichael, Patrick Braga-Henebry, Douglas Thain, Scott J. Emrich |
HPDC | 4 |
| 2010 | A statistical approach to finding overlooked genetic associationsabstractBACKGROUND: Complexity and noise in expression quantitative trait loci (eQTL) studies make it difficult to distinguish potential regulatory relationships among the many interactions. The predominant method of identifying eQTLs finds associations that are significant at a genome-wide level. The vast number of statistical tests carried out on these data make false negatives very likely. Corrections for multiple testing error render genome-wide eQTL techniques unable to detect modest regulatory effects. We propose an alternative method to identify eQTLs that builds on traditional approaches. In contrast to genome-wide techniques, our method determines the significance of an association between an expression trait and a locus with respect to the set of all associations to the expression trait. The use of this specific information facilitates identification of expression traits that have an expression profile that is characterized by a single exceptional association to a locus. Our approach identifies expression traits that have exceptional associations regardless of the genome-wide significance of those associations. This property facilitates the identification of possible false negatives for genome-wide significance. Further, our approach has the property of prioritizing expression traits that are affected by few strong associations. Expression traits identified by this method may warrant additional study because their expression level may be affected by targeting genes near a single locus. RESULTS: We demonstrate our method by identifying eQTL hotspots in Plasmodium falciparum (malaria) and Saccharomyces cerevisiae (yeast). We demonstrate the prioritization of traits with few strong genetic effects through Gene Ontology (GO) analysis of Yeast. Our results are strongly consistent with results gathered using genome-wide methods and identify additional hotspots and eQTLs. CONCLUSIONS: New eQTLs and hotspots found with this method may represent regions of the genome or biological processes that are controlled through few relatively strong genetic interactions. These points of interest warrant experimental investigation. Andrew K. Rider, Geoffrey Siwo, Nitesh V. Chawla, Michael T. Ferdig, Scott J. Emrich |
BMC Bioinform. | 5 |
| 2009 | Harnessing parallelism in multicore clusters with the all-pairs and wavefront abstractionsabstractBoth distributed systems and multicore computers are difficult programming environments. Although the expert programmer may be able to tune distributed and multicore computers to achieve high performance, the non-expert may struggle to achieve a program that even functions correctly. Christopher Moretti, Scott J. Emrich, Kenneth Judd, Douglas Thain |
HPDC | 3 |
| 2007 | Assembling genomes on large-scale parallel computers
Anantharaman Kalyanaraman, Scott J. Emrich, Patrick S. Schnable, Srinivas Aluru |
J. Parallel Distributed Comput. | 2 |
| 2006 | Assembling genomes on large-scale parallel computersabstractAssembly of large genomes from tens of millions of short genomic fragments is computationally demanding requiring hundreds of gigabytes of memory and tens of thousands of CPU hours. New gene-enrichment sequencing strategies are expected to further exacerbate this situation. In this paper, we present a massively parallel genome assembly framework. The unique features of our approach include space-efficient and on-demand algorithms that consume only linear space, and heuristic strategies that reduce the number of expensive pairwise sequence alignments while maintaining assembly quality. As part of the ongoing efforts in maize genome sequencing, we applied our assembly framework to the largest available collection of maize genomic data. We report the partitioning of more than 1.6 million fragments of over 1.25 billion nucleotides total size into genomic islands in 2 hours on 1,024 processors of an IBM BlueGene/L supercomputer. Anantharaman Kalyanaraman, Scott J. Emrich, Patrick S. Schnable, Srinivas Aluru |
IPDPS | 2 |
| 2004 | A comparison of evolved finite state classifiers and interpolated Markov models for improving PCR primer designabstractThis presents results on training both finite state classifiers and interpolated Markov models as classifiers for polymerase chain reaction primers. The goal of the study is to find techniques to decrease the number of primers that fail to amplify correctly within a large genomics project. Standard primer design packages already select primers in a manner consistent with current knowledge of the biophysics of DNA. The classifiers trained in this effort are used to capture lab and organism specific features of primer data and are used to postprocess the output of standard primer design packages. The finite state classifiers in this study are trained with a novel evolutionary algorithm that uses an incremental fitness reward system and multipopulation hybridization. This hybridization is akin to population seeding, not the more usual hybridization of evolutionary computation with other techniques. The interpolated Markov model is a form of Markov model that adapts to data rich and data sparse portions of the training set by using a variable order in its modeling. The interpolated Markov models exhibited slightly superior performance and trains with far higher speed. The finite state classifiers provide a substantially different classification, however, and require less training data. Dan Ashlock, Scott J. Emrich, Kenneth Mark Bryden, Steven M. Corns, Tsui-Jung Wen, Patrick S. Schnable |
CIBCB | 2 |
| 2004 | A strategy for assembling the maize (Zea mays L.) genomeabstractUNLABELLED: Because the bulk of the maize (Zea mays L.) genome consists of repetitive sequences, sequencing efforts are being targeted to its 'gene-rich' fraction. Traditional assembly programs are inadequate for this approach because they are optimized for a uniform sampling of the genome and inherently lack the ability to differentiate highly similar paralogs. RESULTS: We report the development of bioinformatics tools for the accurate assembly of the maize genome. This software, which is based on innovative parallel algorithms to ensure scalability, assembled 730,974 genomic survey sequences fragments in 4 h using 64 Pentium III 1.26 GHz processors of a commodity cluster. Algorithmic innovations are used to reduce the number of pairwise alignments significantly without sacrificing quality. Clone pair information was used to estimate the error rate for improved differentiation of polymorphisms versus sequencing errors. The assembly was also used to evaluate the effectiveness of various filtering strategies and thereby provide information that can be used to focus subsequent sequencing efforts. Scott J. Emrich, Srinivas Aluru, Tsui-Jung Wen, Mahesh Narayanan, Dan Ashlock, Patrick S. Schnable |
Bioinform. | 1 |