Douglas J. Demetrick

dblp:98/7234 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
0since 2021 · last 2011
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 4

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%
Artificial intelligence
1 paper
Kernel, tree and ensemble methods · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › sequence analysis › motif discovery
sequence motif detection
0.112009
Adaptive multi-agent architecture for functional sequence motifs recognition · Bioinform. 2009
Bioinformatics and computational biology › genome annotation
translation initiation site prediction
0.112009
Adaptive multi-agent architecture for functional sequence motifs recognition · Bioinform. 2009
Machine learning › Kernel, tree and ensemble methods › ensemble learning
classifier ensemble
0.012009
Adaptive multi-agent architecture for functional sequence motifs recognition · Bioinform. 2009

Methods — techniques the papers use, named apart from their topics

multi-agent architecture · 0.2genetic algorithm · 0.2classifier ensemble · 0.2
YearPublicationVenuePosition
2011 Employing Machine Learning Techniques for Data Enrichment: Increasing the Number of Samples for Effective Gene Expression Data Analysis
abstract
For certain domains, e.g. bioinformatics, producing more real samples is costly, error prone and time consuming. Therefore, there is a need for an intelligent automated process capable of substituting the real samples by artificial samples that carry the same characteristics as the real samples and hence could be used for running comprehensive testing of new methodologies. Motivated by this need, we describe a novel approach that integrates Probabilistic Boolean Network and genetic algorithm based techniques into a framework that uses some existing real samples as input and successfully produces new samples as output. The new samples will inspire the characteristics of the existing samples without duplicating them. This leads to diversity in the samples and hence a more rich set of samples to be used in testing. The developed framework incorporates two models (perspectives) for sample generation. We illustrate its applicability for producing new gene expression data samples, a high demanding area that has not received attention. The two perspectives employed in the process are based on models that are not closely related, the independence eliminates the bias of having the produced approach covering only certain characteristics of the domain and leading to samples skewed towards one direction. The produced results are very promising in showing the effectiveness, usefulness and applicability of the proposed multi-model framework.
Utku Erdogdu, Mehmet Tan, Reda Alhajj, Faruk Polat, Douglas J. Demetrick, Jon G. Rokne
BIBM5
2009 Adaptive multi-agent architecture for functional sequence motifs recognition
abstract
MOTIVATION: Accurate genome annotation or protein function prediction requires precise recognition of functional sequence motifs. Many computational motif prediction models have been proposed. Due to the complexity of the biological data, it may be desirable to apply an integrated approach that uses multiple models for analysis. RESULTS: In this article, we propose a novel multi-agent architecture for the general purpose of functional sequence motif recognition. The approach takes advantage of the synergy provided by multiple agents through the employment of different agents equipped with distinctive problem solving skills and promotes the collaborations among them through decision maker (DM) agents that work as classifier ensembles. A genetic algorithm-based fusion strategy is applied which offers evolutionary property to the DM agents. The consistency and robustness of the system are maintained by an evolvable agent that mediates the team of the ensemble agents. The combined effort of a recommendation system (Seer) and the self-learning mediator agent yields a successful identification of the most efficient agent deployment scheme at an early stage of the experimentation process, which has the potential of greatly reducing the computational cost of the system. Two concrete systems are constructed that aim at predicting two important sequence motifs-the translational initiation sites (TISs) and the core promoters. With the incorporation of three distinctive problem solver agents, the TIS predictor consistently outperforms most of the state-of-the-art approaches under investigation. Integrating three existing promoter predictors, our system is able to yield consistently good performance. AVAILABILITY: The program (MotifMAS) and the datasets are available upon request.
Reda Alhajj, Douglas J. Demetrick
Bioinform.3
2009 Representative transcript sets for evaluating a translational initiation sites predictor
abstract
BACKGROUND: Translational initiation site (TIS) prediction is a very important and actively studied topic in bioinformatics. In order to complete a comparative analysis, it is desirable to have several benchmark data sets which can be used to test the effectiveness of different algorithms. An ideal benchmark data set should be reliable, representative and readily available. Preferably, proteins encoded by members of the data set should also be representative of the protein population actually expressed in cellular specimens. RESULTS: In this paper, we report a general algorithm for constructing a reliable sequence collection that only includes mRNA sequences whose corresponding protein products present an average profile of the general protein population of a given organism, with respect to three major structural parameters. Four representative transcript collections, each derived from a model organism, have been obtained following the algorithm we propose. Evaluation of these data sets shows that they are reasonable representations of the spectrum of proteins obtained from cellular proteomic studies. Six state-of-the-art predictors have been used to test the usefulness of the construction algorithm that we proposed. Comparative study which reports the predictors' performance on our data set as well as three other existing benchmark collections has demonstrated the actual merits of our data sets as benchmark testing collections. CONCLUSION: The proposed data set construction algorithm has demonstrated its property of being a general and widely applicable scheme. Our comparison with published proteomic studies has shown that the expression of our data set of transcripts generates a polypeptide population that is representative of that obtained from evaluation of biological specimens. Our data set thus represents "real world" transcripts that will allow more accurate evaluation of algorithms dedicated to identification of TISs, as well as other translational regulatory motifs within mRNA sequences. The algorithm proposed by us aims at compiling a redundancy-free data set by removing redundant copies of homologous proteins. The existence of such data sets may be useful for conducting statistical analyses of protein sequence-structure relations. At the current stage, our approach's focus is to obtain an "average" protein data set for any particular organism without posing much selection bias. However, with the three major protein structural parameters deeply integrated into the scheme, it would be a trivial task to extend the current method for obtaining a more selective protein data set, which may facilitate the study of some particular protein structure.
Reda Alhajj, Douglas J. Demetrick
BMC Bioinform.3
2008 Effectiveness of Applying Codon Usage Bias for Translational Initiation Sites Prediction
abstract
The accurate recognition of translational initiation sites (TISs) in genomic, cDNA and mRNA sequences is crucial to identifying the primary structure of the functional gene products - proteins. Many computational methods have been proposed in the literature which apply one or more complicated models that examine a variety of sequence features. In this paper, we propose a novel TIS prediction approach, called codon usage bias agent, which operates solely based on the usage of codon preference; the algorithm only requires O(n) for execution time. The results of the experiments conducted on three benchmark data sets have shown that the proposed approach is very effective and well applicable to solving the problem of TIS recognition.
Reda Alhajj, Douglas J. Demetrick
BIBM3