Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Gail L. Rosen

dblp:81/8353 · also Gail Rosen · DBLP profile ↗
← Back
30ranked-venue papers
4as first author
4since 2021 · last 2025
0000-0003-1763-5750ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 15 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 11 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-authorDatabases, data management, data science and information retrieval · 3Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
5 papers
Bioinformatics and computational biology · 100%
Artificial intelligence
1 paper
Representation and self-supervised learning · 50% Generative modeling · 50%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
metagenomics
1.232025
The Naïve Bayes classifier++ for metagenomic taxonomic classification - query evaluation · Bioinform. 2025
Quikr: a method for rapid reconstruction of bacterial communities via compressive sensing · Bioinform. 2013
NBC: the Naïve Bayes Classification tool webserver for taxonomic classification of metagenomic reads · Bioinform. 2011
Bioinformatics and computational biology › metagenomics
taxonomic classification
1.232025
The Naïve Bayes classifier++ for metagenomic taxonomic classification - query evaluation · Bioinform. 2025
Quikr: a method for rapid reconstruction of bacterial communities via compressive sensing · Bioinform. 2013
NBC: the Naïve Bayes Classification tool webserver for taxonomic classification of metagenomic reads · Bioinform. 2011
Machine learning › Generative modeling › autoregressive model
next-token prediction
0.912025
Context-Aware Regularization with Markovian Integration for Attention-Based Nucleotide Analysis · NeurIPS 2025
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning
0.912025
Context-Aware Regularization with Markovian Integration for Attention-Based Nucleotide Analysis · NeurIPS 2025
Bioinformatics and computational biology › sequence analysis
DNA sequence analysis
0.912025
Context-Aware Regularization with Markovian Integration for Attention-Based Nucleotide Analysis · NeurIPS 2025
Bioinformatics and computational biology › sequence analysis › sequence modeling
genomic sequence modeling
0.912025
Context-Aware Regularization with Markovian Integration for Attention-Based Nucleotide Analysis · NeurIPS 2025
Bioinformatics and computational biology
cancer genomics
0.612022
MetaMutationalSigs: comparison of mutational signature refitting results made easy · Bioinform. 2022
Bioinformatics and computational biology › gene regulation › regulatory element discovery
regulatory element prediction
0.312025
Context-Aware Regularization with Markovian Integration for Attention-Based Nucleotide Analysis · NeurIPS 2025
Bioinformatics and computational biology › metagenomics
taxonomic profiling
0.312025
The Naïve Bayes classifier++ for metagenomic taxonomic classification - query evaluation · Bioinform. 2025
Information theory › signal processing
compressed sensing
0.012013
Quikr: a method for rapid reconstruction of bacterial communities via compressive sensing · Bioinform. 2013

Methods — techniques the papers use, named apart from their topics

transition-matrix loss · 1.7transformer · 1.7n-gram statistics · 1.7naive bayes · 0.9k-mer analysis · 0.9canonical k-mer storage · 0.9visualization · 0.6result aggregation · 0.6k-mer based reconstruction · 0.3compressive sensing · 0.3
YearPublicationVenuePosition
2025 Context-Aware Regularization with Markovian Integration for Attention-Based Nucleotide Analysis
abstract
Transformers have revolutionized nucleotide sequence analysis, yet capturing long‑range dependencies remains challenging. Recent studies show that autoregressive transformers often exhibit Markovian behavior by relying on fixed-length context windows for next-token prediction. However, standard self-attention mechanisms are computationally inefficient for long sequences due to their quadratic complexity and do not explicitly enforce global transition consistency. We introduce CARMANIA (Context-Aware Regularization with Markovian Integration for Attention-Based Nucleotide Analysis), a self-supervised pretraining framework that augments next-token (NT) prediction with a transition-matrix (TM) loss. The TM loss aligns predicted token transitions with empirically derived n-gram statistics from each input sequence, encouraging the model to capture higher-order dependencies beyond local context. This integration enables CARMANIA to learn organism-specific sequence structures that reflect both evolutionary constraints and functional organization. We evaluate CARMANIA across diverse genomic tasks, including regulatory element prediction, functional gene classification, taxonomic inference, antimicrobial resistance detection, and biosynthetic gene cluster classification. CARMANIA outperforms the previous best long-context model by at least 7\%, matches state-of-the-art on shorter sequences (exceeding prior results on 20/40 tasks while running $\sim$2.5$\times$ faster), and shows particularly strong improvements on enhancer and housekeeping gene classification tasks—including up to a 34\% absolute gain in Matthews correlation coefficient (MCC) for enhancer prediction. The TM loss boosts accuracy in 33 of 40 tasks, especially where local motifs or regulatory patterns drive prediction. This enables more effective modeling of sequence-dependent biological features while maintaining robustness across non-coding and low-signal regions. Code available at https://github.com/EESI/carmania.
Mohammadsaleh Refahi, Mahdi Abavisani, Bahrad A. Sokhansanj, James R. Brown, Gail L. Rosen
NeurIPS5
2025 The Naïve Bayes classifier++ for metagenomic taxonomic classification - query evaluation
abstract
MOTIVATION: This study examines the query performance of the NBC++ (Incremental Naive Bayes Classifier) program for variations in canonicality, k-mer size, databases, and input sample data size. We demonstrate that both NBC++ and Kraken2 are influenced by database depth, with macro measures improving as depth increases. However, fully capturing the diversity of life, especially viruses, remains a challenge. RESULTS: NBC++ can competitively profile the superkingdom content of metagenomic samples using a small training database. NBC++ spends less time training and can use a fraction of the memory than Kraken2 but at the cost of long querying time. Major NBC++ enhancements include accommodating canonical k-mer storage (leading to significant storage savings) and adaptable and optimized memory allocation that accelerates query analysis and enables the software to be run on nearly any system. Additionally, the output now includes log-likelihood values for each training genome, providing users with valuable confidence information. AVAILABILITY AND IMPLEMENTATION: Source code and Dockerfile are available at http://github.com/EESI/Naive_Bayes.
Haozhe Duan, Gavin L. A. Hearne, Robi Polikar, Gail L. Rosen
Bioinform.4
2022 MetaMutationalSigs: comparison of mutational signature refitting results made easy
abstract
MOTIVATION: The analysis of mutational signatures is becoming increasingly common in cancer genetics, with emerging implications in cancer evolution, classification, treatment decision and prognosis. Recently, several packages have been developed for mutational signature analysis, with each using different methodology and yielding significantly different results. Because of the non-trivial differences in tools' refitting results, researchers may desire to survey and compare the available tools, in order to objectively evaluate the results for their specific research question, such as which mutational signatures are prevalent in different cancer types. RESULTS: Due to the need for effective comparison of refitting mutational signatures, we introduce a user-friendly software that can aggregate and visually present results from different refitting packages. AVAILABILITY AND IMPLEMENTATION: MetaMutationalSigs is implemented using R and python and is available for installation using Docker and available at: https://github.com/EESI/MetaMutationalSigs.
Palash Pandey, Sanjeevani Arora, Gail L. Rosen
Bioinform.3
2021 Learning, visualizing and exploring 16S rRNA structure using an attention-based deep neural network
abstract
Recurrent neural networks with memory and attention mechanisms are widely used in natural language processing because they can capture short and long term sequential information for diverse tasks. We propose an integrated deep learning model for microbial DNA sequence data, which exploits convolutional neural networks, recurrent neural networks, and attention mechanisms to predict taxonomic classifications and sample-associated attributes, such as the relationship between the microbiome and host phenotype, on the read/sequence level. In this paper, we develop this novel deep learning approach and evaluate its application to amplicon sequences. We apply our approach to short DNA reads and full sequences of 16S ribosomal RNA (rRNA) marker genes, which identify the heterogeneity of a microbial community sample. We demonstrate that our implementation of a novel attention-based deep network architecture, Read2Pheno, achieves read-level phenotypic prediction. Training Read2Pheno models will encode sequences (reads) into dense, meaningful representations: learned embedded vectors output from the intermediate layer of the network model, which can provide biological insight when visualized. The attention layer of Read2Pheno models can also automatically identify nucleotide regions in reads/sequences which are particularly informative for classification. As such, this novel approach can avoid pre/post-processing and manual interpretation required with conventional approaches to microbiome sequence classification. We further show, as proof-of-concept, that aggregating read-level information can robustly predict microbial community properties, host phenotype, and taxonomic classification, with performance at least comparable to conventional approaches. An implementation of the attention-based deep learning network is available at https://github.com/EESI/sequence_attention (a python package) and https://github.com/EESI/seq2att (a command line tool).
Zhengqiao Zhao, Stephen Woloszynek, Felix Agbavor, Joshua Chang Mell, Bahrad A. Sokhansanj, Gail L. Rosen
PLoS Comput. Biol.6
2020 Keeping up with the genomes: efficient learning of our increasing knowledge of the tree of life
abstract
Abstract Background It is a computational challenge for current metagenomic classifiers to keep up with the pace of training data generated from genome sequencing projects, such as the exponentially-growing NCBI RefSeq bacterial genome database. When new reference sequences are added to training data, statically trained classifiers must be rerun on all data, resulting in a highly inefficient process. The rich literature of “incremental learning” addresses the need to update an existing classifier to accommodate new data without sacrificing much accuracy compared to retraining the classifier with all data. Results We demonstrate how classification improves over time by incrementally training a classifier on progressive RefSeq snapshots and testing it on: (a) all known current genomes (as a ground truth set) and (b) a real experimental metagenomic gut sample. We demonstrate that as a classifier model’s knowledge of genomes grows, classification accuracy increases. The proof-of-concept naïve Bayes implementation, when updated yearly, now runs in 1/4 t h of the non-incremental time with no accuracy loss. Conclusions It is evident that classification improves by having the most current knowledge at its disposal. Therefore, it is of utmost importance to make classifiers computationally tractable to keep up with the data deluge. The incremental learning classifier can be efficiently updated without the cost of reprocessing nor the access to the existing database and therefore save storage as well as computation resources.
Zhengqiao Zhao, Alexandru Cristian, Gail L. Rosen
BMC Bioinform.3
2020 Genetic grouping of SARS-CoV-2 coronavirus sequences using informative subtype markers for pandemic spread visualization
abstract
We propose an efficient framework for genetic subtyping of SARS-CoV-2, the novel coronavirus that causes the COVID-19 pandemic. Efficient viral subtyping enables visualization and modeling of the geographic distribution and temporal dynamics of disease spread. Subtyping thereby advances the development of effective containment strategies and, potentially, therapeutic and vaccine strategies. However, identifying viral subtypes in real-time is challenging: SARS-CoV-2 is a novel virus, and the pandemic is rapidly expanding. Viral subtypes may be difficult to detect due to rapid evolution; founder effects are more significant than selection pressure; and the clustering threshold for subtyping is not standardized. We propose to identify mutational signatures of available SARS-CoV-2 sequences using a population-based approach: an entropy measure followed by frequency analysis. These signatures, Informative Subtype Markers (ISMs), define a compact set of nucleotide sites that characterize the most variable (and thus most informative) positions in the viral genomes sequenced from different individuals. Through ISM compression, we find that certain distant nucleotide variants covary, including non-coding and ORF1ab sites covarying with the D614G spike protein mutation which has become increasingly prevalent as the pandemic has spread. ISMs are also useful for downstream analyses, such as spatiotemporal visualization of viral dynamics. By analyzing sequence data available in the GISAID database, we validate the utility of ISM-based subtyping by comparing spatiotemporal analyses using ISMs to epidemiological studies of viral transmission in Asia, Europe, and the United States. In addition, we show the relationship of ISMs to phylogenetic reconstructions of SARS-CoV-2 evolution, and therefore, ISMs can play an important complementary role to phylogenetic tree-based analysis, such as is done in the Nextstrain project. The developed pipeline dynamically generates ISMs for newly added SARS-CoV-2 sequences and updates the visualization of pandemic spatiotemporal dynamics, and is available on Github at https://github.com/EESI/ISM (Jupyter notebook), https://github.com/EESI/ncov_ism (command line tool) and via an interactive website at https://covid19-ism.coe.drexel.edu/.
Zhengqiao Zhao, Bahrad A. Sokhansanj, Charvi Malhotra, Kitty Zheng, Gail L. Rosen
PLoS Comput. Biol.5
2019 16S rRNA sequence embeddings: Meaningful numeric feature representations of nucleotide sequences that are convenient for downstream analyses
abstract
Advances in high-throughput sequencing have increased the availability of microbiome sequencing data that can be exploited to characterize microbiome community structure in situ. We explore using word and sentence embedding approaches for nucleotide sequences since they may be a suitable numerical representation for downstream machine learning applications (especially deep learning). This work involves first encoding ("embedding") each sequence into a dense, low-dimensional, numeric vector space. Here, we use Skip-Gram word2vec to embed k-mers, obtained from 16S rRNA amplicon surveys, and then leverage an existing sentence embedding technique to embed all sequences belonging to specific body sites or samples. We demonstrate that these representations are meaningful, and hence the embedding space can be exploited as a form of feature extraction for exploratory analysis. We show that sequence embeddings preserve relevant information about the sequencing data such as k-mer context, sequence taxonomy, and sample class. Specifically, the sequence embedding space resolved differences among phyla, as well as differences among genera within the same family. Distances between sequence embeddings had similar qualities to distances between alignment identities, and embedding multiple sequences can be thought of as generating a consensus sequence. In addition, embeddings are versatile features that can be used for many downstream tasks, such as taxonomic and sample classification. Using sample embeddings for body site classification resulted in negligible performance loss compared to using OTU abundance data, and clustering embeddings yielded high fidelity species clusters. Lastly, the k-mer embedding space captured distinct k-mer profiles that mapped to specific regions of the 16S rRNA gene and corresponded with particular body sites. Together, our results show that embedding sequences results in meaningful representations that can be used for exploratory analyses or for downstream machine learning applications that require numeric data. Moreover, because the embeddings are trained in an unsupervised manner, unlabeled data can be embedded and used to bolster supervised machine learning tasks.
Stephen Woloszynek, Zhengqiao Zhao, Jian Chen 0043, Gail L. Rosen
PLoS Comput. Biol.4
2018 Extensions to Online Feature Selection Using Bagging and Boosting
abstract
Feature subset selection can be used to sieve through large volumes of data and discover the most informative subset of variables for a particular learning problem. Yet, due to memory and other resource constraints (e.g., CPU availability), many of the state-of-the-art feature subset selection methods cannot be extended to high dimensional data, or data sets with an extremely large volume of instances. In this brief, we extend online feature selection (OFS), a recently introduced approach that uses partial feature information, by developing an ensemble of online linear models to make predictions. The OFS approach employs a linear model as the base classifier, which allows the $l_{0}$ -norm of the parameter vector to be constrained to perform feature selection leading to sparse linear models. We demonstrate that the proposed ensemble model typically yields a smaller error rate than any single linear model, while maintaining the same level of sparsity and complexity at the time of testing.
Gregory Ditzler, Joseph LaBarck, James Ritchie, Gail L. Rosen, Robi Polikar
IEEE Trans. Neural Networks Learn. Syst.4
2018 A Sequential Learning Approach for Scaling Up Filter-Based Feature Subset Selection
abstract
Increasingly, many machine learning applications are now associated with very large data sets whose sizes were almost unimaginable just a short time ago. As a result, many of the current algorithms cannot handle, or do not scale to, today's extremely large volumes of data. Fortunately, not all features that make up a typical data set carry information that is relevant or useful for prediction, and identifying and removing such irrelevant features can significantly reduce the total data size. The unfortunate dilemma, however, is that some of the current data sets are so large that common feature selection algorithms-whose very goal is to reduce the dimensionality-cannot handle such large data sets, creating a vicious cycle. We describe a sequential learning framework for feature subset selection (SLSS) that can scale with both the number of features and the number of observations. The proposed framework uses multiarm bandit algorithms to sequentially search a subset of variables, and assign a level of importance for each feature. The novel contribution of SLSS is its ability to naturally scale to large data sets, evaluate such data in a very small amount of time, and be performed independently of the optimization of any classifier to reduce unnecessary complexity. We demonstrate the capabilities of SLSS on synthetic and real-world data sets.
Gregory Ditzler, Robi Polikar, Gail L. Rosen
IEEE Trans. Neural Networks Learn. Syst.3
2017 Incremental Author Name Disambiguation for Scientific Citation Data
abstract
Name disambiguation is a perennial challenge for any large and growing dataset but is particularly significant for scientific publication data where documents and ideas are linked through citations and depend on highly accurate authorship. Differentiating personal names in scientific publications is a substantial problem as many names are not sufficiently distinct due to the large number of researchers active in most academic disciplines today. As more and more documents and citations are published every year, any system built on this data must be continually retrained and reclassified to remain relevant and helpful. Recently, some incremental learning solutions have been proposed, but most of these have been limited to small-scale simulations and do not exhibit the full heterogeneity of the millions of authors and papers in real world data. In our work, we propose a probabilistic model that simultaneously uses a rich set of metadata and reduces the amount of pairwise comparisons needed for new articles. We suggest an approach to disambiguation that classifies in an incremental fashion to alleviate the need for retraining the model and re-clustering all papers and uses fewer parameters than other algorithms. Using a published dataset, we obtained the highest K-measure which is a geometric mean of cluster and author-class purity. Moreover, on a difficult author block from the Clarivate Analytics Web of Science, we obtain higher precision than other algorithms.
Zhengqiao Zhao, Jason Rollins, Linge Bai, Gail L. Rosen
DSAA4
2015 Fizzy: feature subset selection for metagenomics
abstract
BACKGROUND: Some of the current software tools for comparative metagenomics provide ecologists with the ability to investigate and explore bacterial communities using α- & β-diversity. Feature subset selection--a sub-field of machine learning--can also provide a unique insight into the differences between metagenomic or 16S phenotypes. In particular, feature subset selection methods can obtain the operational taxonomic units (OTUs), or functional features, that have a high-level of influence on the condition being studied. For example, in a previous study we have used information-theoretic feature selection to understand the differences between protein family abundances that best discriminate between age groups in the human gut microbiome. RESULTS: We have developed a new Python command line tool, which is compatible with the widely adopted BIOM format, for microbial ecologists that implements information-theoretic subset selection methods for biological data formats. We demonstrate the software tools capabilities on publicly available datasets. CONCLUSIONS: We have made the software implementation of Fizzy available to the public under the GNU GPL license. The standalone implementation can be found at http://github.com/EESI/Fizzy.
Gregory Ditzler, J. Calvin Morrison, Yemin Lan, Gail L. Rosen
BMC Bioinform.4
2015 A Bootstrap Based Neyman-Pearson Test for Identifying Variable Importance
abstract
Selection of most informative features that leads to a small loss on future data are arguably one of the most important steps in classification, data analysis and model selection. Several feature selection (FS) algorithms are available; however, due to noise present in any data set, FS algorithms are typically accompanied by an appropriate cross-validation scheme. In this brief, we propose a statistical hypothesis test derived from the Neyman-Pearson lemma for determining if a feature is statistically relevant. The proposed approach can be applied as a wrapper to any FS algorithm, regardless of the FS criteria used by that algorithm, to determine whether a feature belongs in the relevant set. Perhaps more importantly, this procedure efficiently determines the number of relevant features given an initial starting point. We provide freely available software implementations of the proposed methodology.
Gregory Ditzler, Robi Polikar, Gail L. Rosen
IEEE Trans. Neural Networks Learn. Syst.3
2014 Scaling a neyman-pearson subset selection approach via heuristics for mining massive data
abstract
Feature subset selection is an important step towards producing a classifier that relies only on relevant features, while keeping the computational complexity of the classifier low. Feature selection is also used in making inferences on the importance of attributes, even when classification is not the ultimate goal. For example, in bioinformatics and genomics feature subset selection is used to make inferences between the variables that best describe multiple populations. Unfortunately, many feature selection algorithms require the subset size to be specified a priori, but knowing how many variables to select is typically a nontrivial task. Other approaches rely on a specific variable subset selection framework to be used. In this work, we examine an approach to feature subset selection works with a generic variable selection algorithm, and our approach provides statistical inference on the number of features that are relevant, which may be unknown to the generic variable selection algorithm. This work extends our previous implementation of a Neyman-Pearson feature selection (NPFS) hypothesis test, which acts as a meta-subset selection algorithm. Specifically, we examine the conservativeness of the NPFS approach by biasing the hypothesis test, and examine other heuristics for NPFS. We include results from carefully designed synthetic datasets. Furthermore, we demonstrate the NPFS's ability to perform on data of a massive scale.
Gregory Ditzler, Matthew Austen, Gail L. Rosen, Robi Polikar
CIDM3
2014 Domain adaptation bounds for multiple expert systems under concept drift
abstract
The ability to learn incrementally from streaming data - either in an online or batch setting - is of crucial importance for a prediction algorithm to learn from environments that generate vast amounts of data, where it is impractical or simply unfeasible to store all historical data. On the other hand, learning from streaming data becomes increasingly difficult when the probability distribution generating the data stream evolves over time, which renders the classification model generated from previously seen data suboptimal or potentially useless. Ensemble systems that employ multiple classifiers may be used to mitigate this effect, but even in such cases some classifiers (experts) become less knowledgeable for predicting on different domains than others as the distribution drifts. Further complication results when labeled data from a prediction (target) domain is not immediately available; hence, causing prediction on the target domain to yield sub-optimal results. In this work, we provide upper bounds on the loss, which hold with high probability, of a multiple expert system trained in such a nonstationary environment with verification latency. Furthermore, we show why a single model selection strategy can lead to undesirable results when learning in such nonstationary streaming settings. We present our analytical results with experiments on simulated as well as real-world data sets, comparing several different ensemble approaches to a single model.
Gregory Ditzler, Gail L. Rosen, Robi Polikar
IJCNN2
2013 Incremental learning of new classes from unbalanced data
abstract
Multiple classifier systems tend to suffer from outvoting when new concept classes need to be learned incrementally. Out-voting is primarily due to existing classifiers being unable to recognize the new class until there is a sufficient number of new classifiers that can influence the ensemble decision. This problem of learning new classes was explicitly addressed in Learn++.NC, our previous work, where ensemble members dynamically adjust their own weights by consulting with each other based on their individual and collective confidence in classifying each concept class. Learn++.NC works remarkably well for learning new concept classes while requiring few ensemble members to do so. Learn++.NC cannot cope with the class imbalance problem, however, as it was not designed to do so. Yet, class imbalance is a common and important problem in machine learning, made even more challenging in an incremental learning setting. In this paper, we extend Learn++.NC so that it can incrementally learn new concept classes even if their instances are drawn from severely imbalanced class distributions. We show that the proposed algorithm is quite robust compared to other state-of-the-art algorithms.
Gregory Ditzler, Gail L. Rosen, Robi Polikar
IJCNN2
2013 Quikr: a method for rapid reconstruction of bacterial communities via compressive sensing
abstract
MOTIVATION: Many metagenomic studies compare hundreds to thousands of environmental and health-related samples by extracting and sequencing their 16S rRNA amplicons and measuring their similarity using beta-diversity metrics. However, one of the first steps--to classify the operational taxonomic units within the sample--can be a computationally time-consuming task because most methods rely on computing the taxonomic assignment of each individual read out of tens to hundreds of thousands of reads. RESULTS: We introduce Quikr: a QUadratic, K-mer-based, Iterative, Reconstruction method, which computes a vector of taxonomic assignments and their proportions in the sample using an optimization technique motivated from the mathematical theory of compressive sensing. On both simulated and actual biological data, we demonstrate that Quikr typically has less error and is typically orders of magnitude faster than the most commonly used taxonomic assignment technique (the Ribosomal Database Project's Naïve Bayesian Classifier). Furthermore, the technique is shown to be unaffected by the presence of chimeras, thereby allowing for the circumvention of the time-intensive step of chimera filtering. AVAILABILITY: The Quikr computational package (in MATLAB, Octave, Python and C) for the Linux and Mac platforms is available at http://sourceforge.net/projects/quikr/.
David Koslicki, Simon Foucart, Gail L. Rosen
Bioinform.3
2012 Forensic identification with environmental samples
abstract
The field of forensics aims to understand the physical biomarkers that make each person unique. Recently, it has been discovered that one of the traits that makes us unique from one another are the composition of the microbial communities found throughout our bodies. For example, identical twins who share the same set of DNA may have vastly different microbial communities in or on various body sites. It was recently discovered that microbial communities can be exploited for forensic identification by clustering samples from individual's skin and objects that they may have previously touched. Typically, this is done by using basic multi-dimensional scaling analysis using phylogenetic distances. In this work, we circumvent the use of phylogenetic distances by using the raw community abundances, and we present an application of kernels for metagenomic data analysis. In addition, we show that strategic selection of features can improve classification accuracy.
Gregory Ditzler, Gail L. Rosen, Robi Polikar
ICASSP2
2012 Transductive learning algorithms for nonstationary environments
abstract
Many traditional supervised machine learning approaches, either on-line or batch based, assume that data are sampled from a fixed yet unknown source distribution. Most incremental learning algorithms also make the same assumption, even though new data are presented over periods of time. Yet, many real-world problems are characterized by data whose distribution change over time, which implies that a classifier may no longer be reliable on future data, a problem commonly referred to as concept drift or learning in nonstationary environments. The issue is further complicated when the problem requires prediction from data obtained at a future time step, for which the labels are not yet available. In this work, we present a transductive learning methodology that uses probabilistic models to aid in computing ensemble classifier voting weights. Assuming the drift is limited in nature, the proposed approach exploits a probabilistic estimate to determine the class responsibility of components in a Gaussian mixture model (GMM), generated from labeled and unlabeled data. A general error bound is provided based on the ensemble decision, the probabilistic estimate of the GMM, and the true labeling function, which, unfortunately is never actually known.
Gregory Ditzler, Gail L. Rosen, Robi Polikar
IJCNN2
2012 Exploiting the Functional and Taxonomic Structure of Genomic Data by Probabilistic Topic Modeling
abstract
In this paper, we present a method that enable both homology-based approach and composition-based approach to further study the functional core (i.e., microbial core and gene core, correspondingly). In the proposed method, the identification of major functionality groups is achieved by generative topic modeling, which is able to extract useful information from unlabeled data. We first show that generative topic model can be used to model the taxon abundance information obtained by homology-based approach and study the microbial core. The model considers each sample as a “document,” which has a mixture of functional groups, while each functional group (also known as a “latent topic”) is a weight mixture of species. Therefore, estimating the generative topic model for taxon abundance data will uncover the distribution over latent functions (latent topic) in each sample. Second, we show that, generative topic model can also be used to study the genome-level composition of “N-mer” features (DNA subreads obtained by composition-based approaches). The model consider each genome as a mixture of latten genetic patterns (latent topics), while each functional pattern is a weighted mixture of the “N-mer” features, thus the existence of core genomes can be indicated by a set of common N-mer features. After studying the mutual information between latent topics and gene regions, we provide an explanation of the functional roles of uncovered latten genetic patterns. The experimental results demonstrate the effectiveness of proposed method.
Xin Chen 0041, Xiaohua Hu 0001, Tze Yee Lim, Xiajiong Shen, E. K. Park, Gail L. Rosen
IEEE ACM Trans. Comput. Biol. Bioinform.6
2011 NBC: the Naïve Bayes Classification tool webserver for taxonomic classification of metagenomic reads
abstract
MOTIVATION: Datasets from high-throughput sequencing technologies have yielded a vast amount of data about organisms in environmental samples. Yet, it is still a challenge to assess the exact organism content in these samples because the task of taxonomic classification is too computationally complex to annotate all reads in a dataset. An easy-to-use webserver is needed to process these reads. While many methods exist, only a few are publicly available on webservers, and out of those, most do not annotate all reads. RESULTS: We introduce a webserver that implements the naïve Bayes classifier (NBC) to classify all metagenomic reads to their best taxonomic match. Results indicate that NBC can assign next-generation sequencing reads to their taxonomic classification and can find significant populations of genera that other classifiers may miss. AVAILABILITY: Publicly available at: http://nbc.ece.drexel.edu.
Gail L. Rosen, Erin R. Reichenberger, Aaron M. Rosenfeld
Bioinform.1
2011 Combining gene prediction methods to improve metagenomic gene annotation
abstract
BACKGROUND: Traditional gene annotation methods rely on characteristics that may not be available in short reads generated from next generation technology, resulting in suboptimal performance for metagenomic (environmental) samples. Therefore, in recent years, new programs have been developed that optimize performance on short reads. In this work, we benchmark three metagenomic gene prediction programs and combine their predictions to improve metagenomic read gene annotation. RESULTS: We not only analyze the programs' performance at different read-lengths like similar studies, but also separate different types of reads, including intra- and intergenic regions, for analysis. The main deficiencies are in the algorithms' ability to predict non-coding regions and gene edges, resulting in more false-positives and false-negatives than desired. In fact, the specificities of the algorithms are notably worse than the sensitivities. By combining the programs' predictions, we show significant improvement in specificity at minimal cost to sensitivity, resulting in 4% improvement in accuracy for 100 bp reads with ~1% improvement in accuracy for 200 bp reads and above. To correctly annotate the start and stop of the genes, we find that a consensus of all the predictors performs best for shorter read lengths while a unanimous agreement is better for longer read lengths, boosting annotation accuracy by 1-8%. We also demonstrate use of the classifier combinations on a real dataset. CONCLUSIONS: To optimize the performance for both prediction and annotation accuracies, we conclude that the consensus of all methods (or a majority vote) is the best for reads 400 bp and shorter, while using the intersection of GeneMark and Orphelia predictions is the best for reads 500 bp and longer. We demonstrate that most methods predict over 80% coding (including partially coding) reads on a real human gut sample sequenced by Illumina technology.
Non Yok, Gail L. Rosen
BMC Bioinform.2
2010 The Effect of Sequence Error and Partial Training Data on BLAST Accuracy
abstract
Metagenomics is the study of environmental samples. Because few tools exist for metagenomic analysis, a natural step has been to utilize the popular homology tool, BLAST, to search for sequence similarity between DNA reads and an administered database. Most biologists use this method today without knowing BLAST's accuracy, especially when a particular taxonomic class is under-represented in the database. The aim of this paper is to benchmark the performance of BLAST for taxonomic classification of metagenomic datasets in a supervised setting, meaning that the database contains microbes of the same class as the `unknown' query DNA reads. We examine well- and under-represented genera and phyla in order to study their effect on the accuracy of BLAST. We investigate the degradation in BLAST accuracy when genome coverage is reduced in the training database as well as the performance when errors are introduced into the query DNA reads. We conclude that on fine-resolution classes, such as genera, the accuracy of BLAST does not degrade very much with under-representation, but in a highly variant class, such as phyla, performance degrades significantly when whole genomes are used in the training database. BLAST accuracy at the genus level is affected greater than phyla when coverage in the training database is reduced or when 1% sequence error is introduced into the query DNA reads. Our analysis includes five-fold cross validation to substantiate our findings.
Steven D. Essinger, Gail L. Rosen
BIBE2
2010 Comparison of Gene Prediction Programs for Metagenomic Data
abstract
This manuscript presents the most rigorous benchmarking of gene annotation algorithms for metagenomic datasets to date. We compare three different programs: GeneMark, MetaGeneAnnotator (MGA) and Orphelia. The comparisons are based on their performances over simulated fragments from hundred species of diverse lineages. We defined three different types of fragments: one type from the intra-coding region and the other types are from the gene edges. The general observation was that performances of all these programs improve as we increase the length of the fragment. On the other hand, intra-coding fragments of our data show a low annotation error in all of the programs if compared to the genes edges.
Non Yok, Gail L. Rosen
BIBE2
2010 Probabilistic topic modeling for genomic data interpretation
abstract
Recently, the concept of a species containing both core and distributed genes, known as the supra- or pangenome theory, has been introduced. In this paper, we aim to develop a new method that is able to analyze the genome-level composition of DNA sequences, in order to characterize a set of common genomic features shared by the same species and tell their functional roles. To achieve this end, we firstly apply a composition-based approach to break down DNA sequences into sub-reads called the `N-mer' and represent the sequences by N-mer frequencies. Then, we introduce the Latent Dirichlet Allocation (LDA) model to study the genome-level statistic patterns (a.k.a. latent topics) of the `N-mer' features. Each estimated latent topic represents a certain component of the whole genome. With the help of the BioJava toolkit, we access to the gene region information of reference sequences from the NCBI database. We use our data mining framework to investigate two areas: 1) do strains within species share similar core and distributed topics? and 2) do genes with similar functional roles contain similar latent topics? After studying the mutual information between latent topics and gene regions, we provide examples of each, where the BioCyc database is used to correlate pathway and reaction information to the genes. The examples demonstrate the effectiveness of proposed method.
Xin Chen 0041, Xiaohua Hu 0001, Xiajiong Shen, Gail L. Rosen
BIBM4
2010 A probabilistic topic-connection model for automatic image annotation
abstract
The explosive increase of image data on Internet has made it an important, yet very challenging task to index and automatically annotate image data. To achieve that end, sophisticated algorithms and models have been proposed to study the correlation between image content and corresponding text description. Despite the success of previous works, however, researchers are still facing two major difficulties that may undermine their effort of providing reliable and accurate annotations for images. The first difficulty is lacking of comprehensive benchmark image dataset with high quality text descriptions. The second difficulty is lacking of effective way to represent the image content and make it associate with the text descriptions. In our paper, we aim to deal with both problems. To deal with the first problem, we utilize Wikipedia as external knowledge source and enrich the ontology structure of ImageNet database with comprehensive and highly-reliable text descriptions from Wikipedia articles. To address the second problem, we develop a Probabilistic Topic-Connection (PTC) model to represent the connection between latent semantic topic in text description and latent patterns from image feature space. We compare the performance of our model with the currently popular Correspondence LDA (Corr-LDA) model under the same automatic image annotation scenario using cross-validation. Experimental results demonstrate that our model is able to well represent the connection between latent semantic topics and latent patterns in image feature space, thus facilitates knowledge organization and understanding of both image and text descriptions.
Xin Chen 0041, Xiaohua Hu 0001, Zhongna Zhou, Caimei Lu, Gail L. Rosen, Tingting He 0003, E. K. Park
CIKM5
2010 Neural network-based taxonomic clustering for metagenomics
abstract
Metagenomic studies inherently involve sampling genetic information from an environment potentially containing thousands of distinctly different microbial organisms. This genetic information is sequenced producing many short fragments (<;500 base pair (bp)); each is tentatively a small representative of the DNA coding structure. Any of the fragments may belong to any of the organisms in the sample, but the relationship is unknown a priori. Furthermore, most of these organisms have not been identified and correspondingly are not represented in any of the publicly available search databases. Our goal is to be able to predict the taxonomic classification of an organism based on the fragments obtained from an environmental sample that may include many (some previously unidentified) organisms. To elucidate the diversity and composition of the sample, we first use a supervised naive Bayes classifier to score the fragments of known genomes, followed by an unsupervised clustering to group fragments from similar organisms together. We are then free to analyze each cluster separately. This is challenging since we are not interested in similar sequences, but sequences that come from similar genomes, which are known to vary widely intra-genomically. Our dataset comprises of an extremely challenging scenario involving clustering fragments at the phyla level, where none of the phyla have been previously seen or identified. We present two variations of our proposed approach, one based on ART and K-means. We show that ART can cluster 500bp fragments from 17 novel phyla at an overall isolation/grouping that is 10% better than K-means and nearly 7 times over chance.
Steven D. Essinger, Robi Polikar, Gail L. Rosen
IJCNN3
2008 Validating models of bacterial chemotaxis by simulating the random motility coefficient
abstract
In order to characterize the random walk of E. coli, biologists have studied several parameters, such as the motility speed, run duration, and random motility coefficient. Previously, biologists indicated that the probability distributions of these parameters may vary depending on the presence or absence of a chemical gradient in the environment. For instance, in a gradient, the cell of E. coli exhibits a biased-random walk. Although it is suggested that the parameter distributions change from unbiased to biased conditions, there are contradicting reports of the actual distributions since they are usually derived from observations of cell movement. In this paper, we consider the problem conversely. We try hypotheses for the parameter distributions respectively for the unbiased and biased cases to simulate random walks. Then we can validate our chemotaxis model through the simulated random motility coefficient, under unbiased and biased environments.
MinJun Kim 0001, Gail L. Rosen
BIBE3
2006 Chemical Source Localization in Unknown Turbulence Using the Cross-Correlation Method
abstract
Estimating the direction of a diffusive source is a difficult problem, and little has been tried to estimate a chemical source subject to turbulence. Turbulence must be addressed if chemical localizer systems are to be effective. We look at how to quantify turbulence and develop a measure to locate a source in two different types of turbulence, modulated and unmodulated plumes. We show that a plume can be modeled linearly on a small-scale and that a wind measure for a stationary sensor array can indicate the direction of a chemical source with reasonable accuracy and time. This measure can be easily implemented in low-power computational electronics and applied to the detection of chemical leaks and illegal substances
Gail L. Rosen, Paul E. Hasler
ICASSP (3)1
2003 Investigation of coding structure in DNA
abstract
We have all heard the term "cracking the genomic code", but is DNA a code in the information theoretic sense? The coined term "genetic code" maps nucleotide triplets (codons) to amino acids. However, this is in a computer coding sense because a codon instruction is performed to output an amino acid sequence. We examine methods to detect redundant coding structures in DNA. First, a finite field framework for a nucleotide symbolic sequence is presented; then approaches to finding the sequence structure associated with error correcting codes are examined. We compare a previously proposed parity-check vector search method to a novel subspace partitioning algorithm. The subspace partitioning algorithm is a general approach to finding any linear coding redundancy. Our method provides an easy way of visualizing coding potential in DNA sequences as shown from the test data.
Gail L. Rosen, Jeffrey D. Moore
ICASSP (2)1
2002 Automatic loudspeaker directivity control for sound field reconstruction
abstract
An enhancement to sound field reconstruction is proposed which improves subjective spatial perception and widens the sweet spot size of a 5 channel surround system. With the popularization of multi-channel audio systems, consumers have access to artificially generated surround sound found in movie audio tracks in addition to recordings featuring the preservation of the original audio space. How to achieve the latter is of question. A recording technique for perceptual sound field reconstruction(PSR) has been proposed by Jim Johnston et al., at AT&T. We now present an enhancement to PSR that can be applied to other surround sound systems. With most commercial surround sound reproduction, sound is directly radiated from all loudspeakers whether the original signal is direct or diffuse/reverberant. In contrast to an attack, reverberation is a far-field multipath phenomenon and reaches the ear as a sequence of decaying reflections. It is therefore desirable to reproduce recordings by scattering the diffuse field as well as radiating the direct field. In the proposed system, the direct field is radiated with a conventional loudspeaker while the diffuse sound field is separated and radiated with a loudspeaker/diffusor combination. The presented design utilizes the ear's ability to localize sound sources while preserving the perceptual spaciousness of the sound field. A digital switch was implemented to split the fields for each direct/diffuse loudspeaker pair.
Gail L. Rosen, Jim Johnston
ICASSP1