EDBT 2026 Demo / reviewers in the wild / expert
Kay C. Wiese
dblp:w/KayCWiese
· DBLP profile ↗
44ranked-venue papers
16as first author
5since 2021 · last 2024
0000-0003-4507-1892ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 28 · 7 first-author · 5 since 2021Artificial intelligence and machine learning · 18 · 11 first-authorHuman-computer interaction and ubiquitous computing · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | AGED-ViT: A Novel Transformer-Based Framework for Diagnosis of Alzheimer's Disease by Leveraging Gene Expression DataabstractAlzheimer’s disease (AD) is a growing global health concern and correct diagnosis is crucial for effective treatment. In this study, we present a novel method for the detection of AD using gene expression data from blood samples. We normalized and combined four publicly available Alzheimer’s datasets and trained a Vision Transformer (ViT) model. This combined dataset had almost seven times more features than patient samples which can cause models to overfit on the training data. To overcome this issue, we employed Linear Discriminant Analysis (LDA) to reduce the dimensionality of the data and noise injection to encourage generalizability and robustness. We then compared our model to several state-of-the-art models that used Support Vector Machines (SVMs), Convolutional Neural Networks (CNNs), and Deep Neural Networks (DNNs). Our model, AGED-ViT, achieved an average accuracy of 88.4% and area under the curve (AUC) of 0.951 on the combined dataset, outperforming previous methods. Our results demonstrate the importance of preprocessing techniques for data with more features than samples to reduce overfitting, as well as the powerful predictive capabilities of ViTs, establishing a foundation for further exploration and optimization of the transformer architecture in the context of genomic diagnosis. This study can contribute to improving the accuracy of AD diagnosis, thus facilitating intervention and leading to a more promising outcome for patients. Albert Guo, Megan Fowler, Kay C. Wiese |
CIBCB | 3 |
| 2023 | A Comparison of Machine Learning Models for Predicting CRISPR/Cas On-target EfficacyabstractCRISPR/Cas (Clustered Regularly Interspaced Short Palindromic Repeats and CRISPR-associated protein) is a powerful technology that can precisely modify DNA, enabling the treatment of several genetic disorders. The CRISPR/Cas system is comprised of a nuclease, which induces the modification, and a sgRNA (single guide RNA), which targets the nuclease to a precise location in the DNA. Designing sgRNAs is time-consuming and resource-intensive, thus computational tools are used to screen sgRNAs for their on-target efficacy. However, to date, models for predicting on-target efficacy have been restricted in complexity due to data limitations. Recently, a large on-target dataset has been published, relieving some of the restraints towards creating larger computational models. Herein, we present a comparison of several types of deep learning models, including CNNs (Convolutional Neural Networks) and RNNs (Recurrent Neural Networks), using the new dataset to predict the on-target efficacy of sgRNAs. After determining which general model performs best on this dataset, we assessed the impact of adding an important biological feature, the GC content, and compared the models to a state-of-the-art competitor, DeepHF [1]. Megan Fowler, Amirhossein Daneshpajouh, Kay C. Wiese |
CIBCB | 3 |
| 2022 | EvoDNN - Evolving Weights, Biases, and Activation Functions in a Deep Neural NetworkabstractClassification, such as classifying cell samples into cancer (malignant) or normal (benign), or the classification of genome regions into functional regions (for example coding regions or promoter regions) are important problems in Computational Biology. For such tasks, we have previously designed an evolutionary deep neural network that in addition to evolving the neuron's weights and biases also evolves (learns) the activation functions for each neuron and called this approach EvoDNN. EvoDNN can employ activation functions that are non-differentiable as it does not rely on back-propagation. This feature is adding flexibility in terms of activation functions EvoDNN can employ. The work presented here extends our previous work on EvoDNN by analyzing a more extensive set of data sets and studying variations of the internal topology of the EvoDNN model. In addition, we study the effect of evolving the weights and biases only while holding the activation function fixed and demonstrate that evolving activation functions indeed provides better performance. We also compare our model to several other popular models and demonstrate superior performance on several data sets. The current code for EvoDNN is available at https://github.com/Payuing/evoDNN. Peiyu Cui, Kay C. Wiese |
CIBCB | 2 |
| 2022 | Continuous chromatin state feature annotation of the human epigenomeabstractMOTIVATION: Segmentation and genome annotation (SAGA) algorithms are widely used to understand genome activity and gene regulation. These methods take as input a set of sequencing-based assays of epigenomic activity, such as ChIP-seq measurements of histone modification and transcription factor binding. They output an annotation of the genome that assigns a chromatin state label to each genomic position. Existing SAGA methods have several limitations caused by the discrete annotation framework: such annotations cannot easily represent varying strengths of genomic elements, and they cannot easily represent combinatorial elements that simultaneously exhibit multiple types of activity. To remedy these limitations, we propose an annotation strategy that instead outputs a vector of chromatin state features at each position rather than a single discrete label. Continuous modeling is common in other fields, such as in topic modeling of text documents. We propose a method, epigenome-ssm-nonneg, that uses a non-negative state space model to efficiently annotate the genome with chromatin state features. We also propose several measures of the quality of a chromatin state feature annotation and we compare the performance of several alternative methods according to these quality measures. RESULTS: We show that chromatin state features from epigenome-ssm-nonneg are more useful for several downstream applications than both continuous and discrete alternatives, including their ability to identify expressed genes and enhancers. Therefore, we expect that these continuous chromatin state features will be valuable reference annotations to be used in visualization and downstream analysis. AVAILABILITY AND IMPLEMENTATION: Source code for epigenome-ssm is available at https://github.com/habibdanesh/epigenome-ssm and Zenodo (DOI: 10.5281/zenodo.6507585). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Habib Daneshpajouh, Neda Shokraneh Kenari, Shohre Masoumi, Kay C. Wiese, Maxwell W. Libbrecht |
Bioinform. | 5 |
| 2022 | SigTools: exploratory visualization for genomic signalsabstractMOTIVATION: With the advancement of sequencing technologies, genomic data sets are constantly being expanded by high volumes of different data types. One recently introduced data type in genomic science is genomic signals, which are usually short-read coverage measurements over the genome. To understand and evaluate the results of such studies, one needs to understand and analyze the characteristics of the input data. RESULTS: SigTools is an R-based genomic signals visualization package developed with two objectives: (i) to facilitate genomic signals exploration in order to uncover insights for later model training, refinement and development by including distribution and autocorrelation plots; (ii) to enable genomic signals interpretation by including correlation and aggregation plots. In addition, our corresponding web application, SigTools-Shiny, extends the accessibility scope of these modules to people who are more comfortable working with graphical user interfaces instead of command-line tools. AVAILABILITY AND IMPLEMENTATION: SigTools source code, installation guide and manual is freely available on http://github.com/shohre73. Shohre Masoumi, Maxwell W. Libbrecht, Kay C. Wiese |
Bioinform. | 3 |
| 2019 | EvoDNN - An Evolutionary Deep Neural Network with Heterogeneous Activation FunctionsabstractMany problems in Computational Biology and Bioinformatics involve classification, such as the classification of cell samples into malignant (cancer) or benign (normal). For such tasks, we propose EvoDNN, an evolutionary deep neural network that employs an evolutionary algorithm to evolve deep heterogeneous feed-forward neural networks. While the majority of current feed-forward neural networks employ user defined homogeneous activation functions, EvoDNN creates heterogeneous multi-layer networks where each neuron's activation function is not statically defined by the user, but dynamically optimized during evolution. The main advantage offered by EvoDNN lies in that the activation functions do not need to be differentiable. This feature gives users a great degree of flexibility over which activation functions EvoDNN can utilize. This paper demonstrates how EvoDNN can simultaneously optimize each neuron's weight, bias, and activation function, and empirically shows a superior performance compared to a backpropagation-trained feed-forward neural network at the cost of additional training time. In addition, advantages of the deep architecture of EvoDNN over our earlier approach, EvoNN, which employed a single hidden layer are discussed. Peiyu Cui, Boris Shabash, Kay C. Wiese |
CEC | 3 |
| 2017 | Numerical integration methods and layout improvements in the context of dynamic RNA visualizationabstractBACKGROUND: RNA visualization software tools have traditionally presented a static visualization of RNA molecules with limited ability for users to interact with the resulting image once it is complete. Only a few tools allowed for dynamic structures. One such tool is jViz.RNA. Currently, jViz.RNA employs a unique method for the creation of the RNA molecule layout by mapping the RNA nucleotides into vertexes in a graph, which we call the detailed graph, and then utilizes a Newtonian mechanics inspired system of forces to calculate a layout for the RNA molecule. The work presented here focuses on improvements to jViz.RNA that allow the drawing of RNA secondary structures according to common drawing conventions, as well as dramatic run-time performance improvements. This is done first by presenting an alternative method for mapping the RNA molecule into a graph, which we call the compressed graph, and then employing advanced numerical integration methods for the compressed graph representation. RESULTS: Comparing the compressed graph and detailed graph implementations, we find that the compressed graph produces results more consistent with RNA drawing conventions. However, we also find that employing the compressed graph method requires a more sophisticated initial layout to produce visualizations that would require minimal user interference. Comparing the two numerical integration methods demonstrates the higher stability of the Backward Euler method, and its resulting ability to handle much larger time steps, a high priority feature for any software which entails user interaction. CONCLUSION: The work in this manuscript presents the preferred use of compressed graphs to detailed ones, as well as the advantages of employing the Backward Euler method over the Forward Euler method. These improvements produce more stable as well as visually aesthetic representations of the RNA secondary structures. The results presented demonstrate that both the compressed graph representation, as well as the Backward Euler integrator, greatly enhance the run-time performance and usability. The newest iteration of jViz.RNA is available at https://jviz.cs.sfu.ca/download/download.html . Boris Shabash, Kay C. Wiese |
BMC Bioinform. | 2 |
| 2017 | RNA Visualization: Relevance and the Current State-of-the-Art Focusing on PseudoknotsabstractRNA visualization is crucial in order to understand the relationship that exists between RNA structure and its function, as well as the development of better RNA structure prediction algorithms. However, in the context of RNA visualization, one key structure remains difficult to visualize: Pseudoknots. Pseudoknots occur in RNA folding when two secondary structural components form base-pairs between them. The three-dimensional nature of these components makes them challenging to visualize in two-dimensional media, such as print media or screens. In this review, we focus on the advancements that have been made in the field of RNA visualization in two-dimensional media in the past two decades. The review aims at presenting all relevant aspects of pseudoknot visualization. We start with an overview of several pseudoknotted structures and their relevance in RNA function. Next, we discuss the theoretical basis for RNA structural topology classification and present RNA classification systems for both pseudoknotted and non-pseudoknotted RNAs. Each description of RNA classification system is followed by a discussion of the software tools and algorithms developed to date to visualize RNA, comparing the different tools' strengths and shortcomings. Boris Shabash, Kay C. Wiese |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2013 | pEvoSAT: a novel permutation based genetic algorithm for solving the boolean satisfiability problemabstractIn this paper we introduce pEvoSAT, a permutation based Genetic Algorithm (GA), designed to solve the boolean satisfiability (SAT) problem when it is presented in the conjunctive normal form (CNF). The use of permutation based representation allows the algorithm to take advantage of domain specific knowledge such as unit propagation, and pruning. In this paper, we explore and characterize the behavior of our algorithm. This paper also presents the comparison of pEvoSAT to GASAT, a leading implementation of GAs for the solving of CNF instances. Boris Shabash, Kay C. Wiese |
GECCO | 2 |
| 2012 | Non-coding RNA gene finding with combined partial covariance modelsabstractCovariance models provide excellent accuracy for ncRNA homology search. However, high computational complexity has limited their usefulness. This research improves the covariance model's search efficiency by building combined models for a group of different RNA families, which is selected using a clustering strategy. A series of combined partial covariance models are built from the stem loop structural elements that the ncRNA gene families share. Experimental results suggest that for most RNA gene families investigated, our combination search method successfully provides run time improvement with acceptable accuracy. Although there still exist limitations such as recall loss for a few RNA gene families, this novel combination approach has implications for future studies of reducing covariance model's search complexity. Wenbo Jiang 0004, Kay C. Wiese |
CIBCB | 2 |
| 2011 | Improving splice-junctions classification employing a novel encoding schema and decision-treeabstractSplice junctions are important regions in genes, which have been studied in many studies in Genetics. Recently some attempts in computer science have been made to use computer power in distinguishing the different splice junctions and non-junctions regions in genes. Ambiguity in the measurements of nucleotides is an important issue in dealing with these regions. In this paper a novel method is proposed along with an encoding schema which take ambiguities into account using probabilistic intuitions. The method is based on Decision Trees, using K Nearest Negihbours and Support Vector Machines. The results have shown the significance of using the proposed encoding schema and classification method. Amin Yazdani Salekdeh, Kay C. Wiese |
IEEE Congress on Evolutionary Computation | 2 |
| 2011 | Secondary structure element voting for RNA gene findingabstractAn exploration of the use of multiple secondary structure elements for structural RNA gene finding is conducted. The secondary structure models are combined through a multilayer voting system which first combines the probability output of support vector machines and then combines the results of those votes to predict whether a sequence is a structural RNA gene or not. It is found that the voting in the first layer of the system has significant impact on the performance of individual secondary structure element models with improvements in classification results of up to 56%. Likewise, gains in classification F-measure over 0.6 were seen when two secondary structure element model predictions were voted together. When all the secondary structure element models were used in voting, an accuracy of over 93% was achieved by the secondary structure RNA gene classification system. Nicholas George Erho, Kay C. Wiese |
CIBCB | 2 |
| 2011 | Combined covariance model for non-coding RNA gene findingabstractThe use of covariance models in finding non-coding RNA gene members in genome sequence databases has been shown quite effective in many studies. However, it has a significant drawback, which is the very large computational burden. A combined covariance model is proposed to reduce the search complexity when a genome sequence is searched for more than one ncRNA gene family. The covariance models that are combined are selected using a hierarchical clustering algorithm. This study shows that when a small number of original covariance models are combined, the combined covariance model can find members from all original ncRNA families thus successfully reducing the search time. Wenbo Jiang 0004, Kay C. Wiese |
CIBCB | 2 |
| 2010 | A study of RNA secondary structure prediction using different mutation operatorsabstractRibonucleic Acid (RNA) has important structural and functional roles in the cell and plays roles in many stages of protein synthesis as well. The functions of RNA molecules are determined largely by their three-dimensional structure. SARNA-Predict has shown excellent results in predicting RNA secondary structure based on Simulated Annealing (SA). SARNA-Predict uses a permutation-based representation to the RNA secondary structure and this paper investigates the impact of the mutation operators in this algorithm. Experiments were performed using a sample of eleven sequences from four RNA classes. The results presented in this paper demonstrate that SARNA-Predict using the percentage swap translocating mutation operator can produce similar results when compared with previous research. Furthermore, the new operator has the potential of reaching a solution with a lower free energy. This supports the use of the proposed operator on RNA secondary structure prediction of other known structures. Herbert H. Tsang, Tiancheng Jiang, Kay C. Wiese, Christian Jacob 0001 |
IEEE Congress on Evolutionary Computation | 3 |
| 2010 | An exploration of individual RNA structural elements in RNA gene findingabstractThis paper explores the use of RNA structural elements for RNA gene finding. A classification experiment is performed in which several support vector machine models, based on the properties of individual RNA secondary structure elements, are trained and tested, revealing the structural elements which have properties useful for RNA gene finding. The study finds that the external loop and structure elements with classification accuracies of over 84% and the stemloop, hairpin, and tail elements with classification accuracies around 70% are the most likely structural elements to be successfully exploited by future RNA gene finders. Nicholas George Erho, Kay C. Wiese |
CIBCB | 2 |
| 2010 | Expanded study of efn2 thermodynamic model performance on RnaPredict, an evolutionary algorithm for RNA foldingabstractThe shape that organic molecules such as biopolymers form within organic systems largely determines the function said molecules perform. RNA is a biopolymer that plays a central part in several stages of protein synthesis, and also has structural, functional, and regulatory roles in the cell. In an ab initio case most common structure prediction techniques employ minimization of the free energy of a given RNA molecule via a thermodynamic model. RnaPredict is an evolutionary algorithm for RNA folding. This paper compares the performance of an advanced thermodynamic model, efn2, against the stacking-energy thermodynamic models INN and INN-HB on a test set containing 24 sequences from 4 rRNA subtypes. The prediction accuracy of efn2 is demonstrated on a majority of test sequences. A comparison is also made with the mfold prediction algorithm which demonstrated RnaPredict's comparable performance. Kay C. Wiese, Andrew Hendriks |
CIBCB | 1 |
| 2010 | SARNA-Predict: Accuracy Improvement of RNA Secondary Structure Prediction Using Permutation-Based Simulated AnnealingabstractRibonucleic acid (RNA), a single-stranded linear molecule, is essential to all biological systems. Different regions of the same RNA strand will fold together via base pair interactions to make intricate secondary and tertiary structures that guide crucial homeostatic processes in living organisms. Since the structure of RNA molecules is the key to their function, algorithms for the prediction of RNA structure are of great value. In this article, we demonstrate the usefulness of SARNA-Predict, an RNA secondary structure prediction algorithm based on Simulated Annealing (SA). A performance evaluation of SARNA-Predict in terms of prediction accuracy is made via comparison with eight state-of-the-art RNA prediction algorithms: mfold, Pseudoknot (pknotsRE), NUPACK, pknotsRG-mfe, Sfold, HotKnots, ILM, and STAR. These algorithms are from three different classes: heuristic, dynamic programming, and statistical sampling techniques. An evaluation for the performance of SARNA-Predict in terms of prediction accuracy was verified with native structures. Experiments on 33 individual known structures from eleven RNA classes (tRNA, viral RNA, antigenomic HDV, telomerase RNA, tmRNA, rRNA, RNaseP, 5S rRNA, Group I intron 23S rRNA, Group I intron 16S rRNA, and 16S rRNA) were performed. The results presented in this paper demonstrate that SARNA-Predict can out-perform other state-of-the-art algorithms in terms of prediction accuracy. Furthermore, there is substantial improvement of prediction accuracy by incorporating a more sophisticated thermodynamic model (efn2). Herbert H. Tsang, Kay C. Wiese |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2009 | RNA pseudoknot prediction via an evolutionary algorithmabstractBeyond its critical role in protein synthesis, RNA has vital structural, functional, and regulatory roles in the cell. The shape of an RNA molecule primarily determines its function in organic systems, so there is notable interest in the computational prediction of RNA structure. Pseudoknots are relatively rare but important structural elements which are difficult to predict computationally. RnaPredict is an evolutionary algorithm (EA) developed for the prediction of RNA secondary structure. This research evaluates RnaPredict after its enhancement with the thermodynamic model from HotKnots, a model specifically designed to compute free energies of structures containing pseudoknots. The performance of the EA is evaluated against the original HotKnots algorithm. RnaPredict significantly improved upon the sensitivity and specificity of structures predicted by HotKnots. Kay C. Wiese, Andrew Hendriks |
IEEE Congress on Evolutionary Computation | 1 |
| 2009 | Impact of an enhanced thermodynamic model on RnaPredict, an evolutionary algorithm for RNA secondary structure predictionabstractRNA has important structural, functional, and regulatory parts in the cell as well as a critical role in multiple stages of protein synthesis. An RNA molecule's shape largely determines its function in an organic system. Accordingly, computational RNA structural prediction methods are of significant interest. For ab initio cases where only an RNA sequence is known, structure prediction techniques typically employ free energy minimization of a given RNA molecule via a thermodynamic model. Unfortunately, the minimum free energy structure is rarely the native structure. This is thought to be due to errors in the experimentally determined thermodynamic model parameters. RnaPredict is an evolutionary algorithm designed for the prediction of RNA secondary structure; it currently utilizes the stacking-energy thermodynamic models INN and INN-HB. The effect of an enhanced model, efn2, on RnaPredict is investigated. The efn2 model significantly improved the sensitivity and specificity of the majority of structures evaluated. Kay C. Wiese, Andrew Hendriks |
IEEE Congress on Evolutionary Computation | 1 |
| 2009 | rnaDesign: Local search for RNA secondary structure designabstractThe RNA secondary structure design (SSD) problem is a recently emerging research topic motivated by applications such as customized drug design and the self-assembly of RNA nano-objects. This paper presents a novel local search algorithm, rnaDesign for SSD solving. An evaluation of the algorithm performance in terms of sequence affinity and structure specificity is made through comparison with another algorithm, RNAinverse. Experiments were performed on RNA secondary structures including three biologically existing data sets and one random structure set. Empirical results show that rnaDesign outperforms RNAinverse in terms of structure designability; sequences designed by rnaDesign also exhibit better thermodynamic stability with relatively lower folding energy. Furthermore, we demonstrate through parameter tuning experiments that using a combination of heuristic search strategies leads to better design performance; there also exists a strong correlation between the heuristic values in use and solution quality. Denny C. Dai, Herbert H. Tsang, Kay C. Wiese |
CIBCB | 3 |
| 2009 | Performance prediction for RNA design using parametric and non-parametric regression modelsabstractEmpirical algorithm study involves tuning various parameter settings in order to achieve an optimal performance. It is also experimentally known that algorithm performance varies across problem instances. In stochastic local search (metaheuristics) paradigm, search efficiency is correlated to the empirical hardness of the underlying combinatorial optimization problem itself. Therefore, investigating these correlations are of crucial importance towards the design of robust algorithmic solutions. To achieve this goal, an accurate prediction of algorithm performance is a prerequisite, since it allows an automatic tuning of parameter settings on a per-problem base. In this work, we investigate using parametric & non-parametric regression models for algorithm performance prediction for the RNA Secondary Structure Design problem (SSD). Empirical results show our non-parametric methods achieve a higher prediction accuracy on biologically existing data, where biological data exhibits a higher degree of local similarity among individual instances. We also found that using a non-parametric regression tree model (CART) provides insight into studying the empirical hardness of solving the SSD problem. Denny C. Dai, Kay C. Wiese |
CIBCB | 2 |
| 2009 | SARNA-ensemble-predict: The effect of different dissimilarity metrics on a novel ensemble-based RNA secondary structure prediction algorithmabstractRecently, there is a resurgence of interest in the RNA secondary structure prediction problem due to the discovery of many new families of non-coding RNAs with a variety of functions. This paper describes and presents a novel algorithm for RNA secondary structure prediction based on an ensemble-based approach. An evaluation of the performance in terms of sensitivity and specificity is made. Experiments were performed on eleven structures from four RNA classes (RNaseP, Group I intron 16S rRNA, Group I intron 23S rRNA and 16S rRNA). Three RNA secondary structure similarity metrics (base pair distance, tree edit distance, and thermodynamic energy distance) and their effects on the clustering algorithm were explored. The significant contribution of this paper is in the examining of the various results from employing different dissimilarity metrics. Overall, the base pair distance dissimilarity metric shows better results with the other two distance metrics (tree edit distance and thermodynamic energy distance). The results presented in this paper demonstrate that SARNA-Ensemble-Predict can give comparable performance to a state-of-the-art algorithm Sfold in terms of sensitivity. Herbert H. Tsang, Kay C. Wiese |
CIBCB | 2 |
| 2008 | SARNA-Predict-pk: Predicting RNA secondary structures including pseudoknotsabstractPseudoknots are RNA tertiary structures which perform essential biological functions. This paper presents SARNA-Predict-pk, an algorithm for pseudoknotted RNA secondary structure prediction based on Simulated Annealing (SA). The research presented here extends previous work of SARNA-Predict and incorporates a new thermodynamic model into the algorithm, effectively enabling it to predict pseudo-knotted RNA structures. An evaluation of the performance of SARNA-Predict-pk in terms of prediction accuracy is made via comparison with several state-of-the-art prediction algorithms. We measured the sensitivity and specificity using five prediction algorithms. Three of these are dynamic programming algorithms: Pseudoknot (pknotsRE), NUPACK, and pknotsRG-mfe. The other two are heuristic algorithms: SARNA-Predict-pk and HotKnots algorithms. An evaluation for the performance of SARNA-Predict-pk in terms of prediction accuracy was verified with native structures. Experiments on ten individual known structures from six RNA classes (tRNA, viral RNA, anti-genomic HDV, telomerase RNA, tmRNA, and RNaseP) were performed. The results presented in this paper demonstrate that SARNA-Predict-pk can out-perform other state-of-the-art algorithms in terms of prediction accuracy. Herbert H. Tsang, Kay C. Wiese |
CIBCB | 2 |
| 2008 | A hybrid clustering/evolutionary algorithm for RNA foldingabstractRNA is central in several stages of protein synthesis, and also has structural, functional, and regulatory roles in the cell. The shape of organic molecules such as RNA largely determines their function within an organic system, thus methods for the computational prediction of structure are sought after. In the ab initio case where only the RNA sequence is known, the currently dominant structure prediction techniques employ minimization of the free energy of a given RNA molecule via a thermodynamic model. However, the minimum free energy structure is rarely the native structure; this is thought to be due to errors in the thermodynamic model parameters, which are experimentally determined. Cluster analysis performed by [6] on a sampling of structures from a Boltzmann weighted ensemble determined that the best cluster centroid had an improved sensitivity and significantly improved positive predictive value over the minimum free energy structure in the ensemble. Based on this result, we investigated the combination of an existing evolutionary algorithm for RNA secondary structure prediction with a clustering algorithm. Kay C. Wiese, Andrew Hendriks |
CIBCB | 1 |
| 2008 | RnaPredict-An Evolutionary Algorithm for RNA Secondary Structure PredictionabstractThis paper presents two in-depth studies on RnaPredict, an evolutionary algorithm for RNA secondary structure prediction. The first study is an analysis of the performance of two thermodynamic models, Individual Nearest Neighbor (INN) and Individual Nearest Neighbor Hydrogen Bond (INN-HB). The correlation between the free energy of predicted structures and the sensitivity is analyzed for 19 RNA sequences. Although some variance is shown, there is a clear trend between a lower free energy and an increase in true positive base pairs. With increasing sequence length, this correlation generally decreases. In the second experiment, the accuracy of the predicted structures for these 19 sequences are compared against the accuracy of the structures generated by the mfold dynamic programming algorithm (DPA) and also to known structures. RnaPredict is shown to outperform the minimum free energy structures produced by mfold and has comparable performance when compared to sub-optimal structures produced by mfold. Kay C. Wiese, Alain Deschênes, Andrew Hendriks |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2007 | The significance of thermodynamic models in the accuracy improvement of RNA secondary structure prediction using permutation-based simulated annealingabstractRibonucleic acid, a single stranded linear molecule, is essential to all biological systems. Different regions of the same RNA strand will fold together via base pair interactions to make intricate secondary and tertiary structures that guide crucial homeostatic processes in living organisms. Since the structure of RNA molecules is key to their function, algorithms for the prediction of RNA structure are of great value. This paper discusses significant improvements made toSARNA-Predict, an RNA secondary structure prediction algorithm based on Simulated Annealing (SA). One major improvement is the incorporation of a sophisticated thermodynamic model (efn2). This model is used bymfoldto rank sub-optimal structures, but cannot be used directly bymfoldduring the structure prediction. Experiments on eight individual known structures from four RNA classes (5S rRNA, Group I intron 23S rRNA, Group I intron 16S rRNA and 16S rRNA) were performed. The data demonstrate the robustness and the effectiveness of our improved prediction algorithm. The new algorithm shows results which surpass the dynamic programming algorithmmfoldin terms of prediction accuracy on all tested structures. Herbert H. Tsang, Kay C. Wiese |
IEEE Congress on Evolutionary Computation | 2 |
| 2007 | SARNA-Predict: A Study of RNA Secondary Structure Prediction Using Different Annealing SchedulesabstractThis paper presents an algorithm for RNA secondary structure prediction based on simulated annealing (SA) and also studies the effect of using different types of annealing schedules. SA is known to be effective in solving many different types of minimization problems and for being able to approximate global minima in the solution space. Based on free energy minimization techniques, this permutation-based SA algorithm heuristically searches for the structure with a free energy value close to the minimum free energy DeltaG for that strand, within given constraints. Other contributions of this paper include the use of permutation-based encoding for RNA secondary structure and the swap mutation operator. Also, a detailed study of the convergence behavior of the algorithm is conducted and various annealing schedules are investigated. An evaluation of the performance of the new algorithm in terms of prediction accuracy is made via comparison with the dynamic programming algorithm mfold for thirteen individual known structures from four RNA classes (5S rRNA, Group I intron 23 rRNA, Group I intron 16S rRNA and 16S rRNA). Although dynamic programming algorithms for RNA folding are guaranteed to give the mathematically optimal (minimum energy) structure, the fundamental problem of this approach seems to be that the thermodynamic model is only accurate within 5-10%. Therefore, it is difficult for a single sequence folding algorithm to resolve which of the plausible lowest-energy structure is correct. The new algorithm showed comparable results with mfold and demonstrated a slightly higher specificity Herbert H. Tsang, Kay C. Wiese |
CIBCB | 2 |
| 2006 | Graph Drawing Tools for Bioinformatics Research: An OverviewabstractMuch of the data in bioinformatics and data relationships can be represented by graphs. General purpose graph drawing tools are available, but not all are adequate for bioinformatics graphs as these tend to grow quite large. The specific purpose for our research was to investigate the applicability of existing graph drawing tools for visualization of very large phylogenetic trees. In this paper we describe some of the general functions and features that are required for exploring and comparing graphs in bioinformatics. We include a description of a selection of existing tools, to give an overview of the current abilities of graph drawing software and to analyze their advantages and drawbacks in the bioinformatics domain. Kay C. Wiese, Christina Eicher |
CBMS | 1 |
| 2006 | jViz.Rna - An Interactive Graphical Tool for Visualizing RNA Secondary Structure Including PseudoknotsabstractIn order for structure prediction researchers to better understand the results of their algorithms and to enable life science researchers to interpret RNA structure easily, it is helpful to provide them with a flexible and powerful tool for RNA secondary structure visualization. jViz.Rna is a multi-platform visualization tool capable of displaying RNA secondary structures encoded in a variety of file formats. A single structure can be shown using the linear Feynman, circular Feynman, dot plot, and classical structure visualization models. The resulting drawings are dynamic and can easily be further modified by the user. Any of the drawings produced can be saved to disk enabling easy dissemination. The unique usage of a spring model for classical structure drawing allows for clear visualization of pseudoknots with minimal overlaps. The addition of a locality tool allows for the isolation of pseudoknotted regions or other regions of interest Kay C. Wiese, Edward Glen |
CBMS | 1 |
| 2006 | A Detailed Analysis of Parallel Speedup in P-RnaPredict - An Evolutionary Algorithm for RNA Secondary Structure PredictionabstractThe function of an RNA molecule is primarily established by its physical shape. As current physical structure determination methods are time consuming and expensive, there is great interest in finding computational structure prediction methods. P-RnaPredict is a parallel evolutionary algorithm for RNA secondary structure prediction. Two sets of experiments are performed on 5 known structures from 3 RNA classes (5S rRNA, Group I intron 16S rRNA, and 16S rRNA). The first determines the actual speedup, and the second evaluates the performance of P-RnaPredict through comparison to mfold. P-RnaPredict succeeds in predicting structures with higher true positive base pair counts and lower false positives than mfold on specific sequences. Kay C. Wiese, Andrew Hendriks |
IEEE Congress on Evolutionary Computation | 1 |
| 2006 | SARNA-Predict: A Simulated Annealing Algorithm for RNA Secondary Structure PredictionabstractRibonucleic acid (RNA) plays fundamental roles in cellular processes and its structure is directly related to its functions. This paper describes and presents a novel algorithm for RNA secondary structure prediction based on simulated annealing (SA). SA is known to be effective in solving many different types of minimization problems and for finding the global minima in the solution space. Based on free energy minimization techniques, this permutation-based SA algorithm heuristically searches for the structure with a free energy value close to the minimum free energy DeltaG for that strand, within given constraints. A detailed study of the convergence behavior of the algorithm is conducted and various cooling schedules are investigated. An evaluation of the performance of the new algorithm in terms of prediction accuracy is made via comparison with the dynamic programming algorithm mfold for eight individual known structures from three RNA classes (5S rRNA, Group I intron 16S rRNA and 16S rRNA). The significant contribution of this algorithm is in showing comparable results with the most common dynamic programming prediction application mfold and surpassing results from an evolutionary algorithm (EA) Herbert H. Tsang, Kay C. Wiese |
CIBCB | 2 |
| 2006 | Analysis of Thermodynamic Models and Performance in RnaPredict - An Evolutionary Algorithm for RNA FoldingabstractTwo extensive analyzes on RnaPredict, an evolutionary algorithm for RNA folding, are presented here. The first study evaluates the performance of individual nearest neighbor (INN) and individual nearest neighbor-hydrogen bond (INN-HB), two stacking-energy thermodynamic models; the criteria for comparison is the correlation between the prediction accuracy and the free energy of predicted structures for 9 RNA sequences. Despite some variance, a trend between lower free energies and increases in true positive base pairs is apparent. In general, this correlation decreases as the sequence length increases. The second study compares the performance of RnaPredict against the mfold dynamic programming algorithm (DPA) on the same sequences in terms of specificity and sensitivity. The results indicate that RnaPredict has comparable performance to mfold on sub-optimal structures, and outperforms mfold's minimum free energy structures Kay C. Wiese, Andrew Hendriks, Alain Deschênes |
CIBCB | 1 |
| 2006 | Comparison of P-RnaPredict and mfold - algorithms for RNA secondary structure predictionabstractMOTIVATION: Ribonucleic acid is vital in numerous stages of protein synthesis; it also possesses important functional and structural roles within the cell. The function of an RNA molecule within a particular organic system is principally determined by its structure. The current physical methods available for structure determination are time-consuming and expensive. Hence, computational methods for structure prediction are sought after. The energies involved by the formation of secondary structure elements are significantly greater than those of tertiary elements. Therefore, RNA structure prediction focuses on secondary structure. RESULTS: We present P-RnaPredict, a parallel evolutionary algorithm for RNA secondary structure prediction. The speedup provided by parallelization is investigated with five sequences, and a dramatic improvement in speedup is demonstrated, especially with longer sequences. An evaluation of the performance of P-RnaPredict in terms of prediction accuracy is made through comparison with 10 individual known structures from 3 RNA classes (5S rRNA, Group I intron 16S rRNA and 16S rRNA) and the mfold dynamic programming algorithm. P-RnaPredict is able to predict structures with higher true positive base pair counts and lower false positives than mfold on certain sequences. AVAILABILITY: P-RnaPredict is available for non-commercial usage. Interested parties should contact Kay C. Wiese ([email protected]). Kay C. Wiese, Andrew Hendriks |
Bioinform. | 1 |
| 2006 | Ebbie: automated analysis and storage of small RNA cloning data using a dynamic web serverabstractBACKGROUND: DNA sequencing is used ubiquitously: from deciphering genomes to determining the primary sequence of small RNAs (smRNAs). The cloning of smRNAs is currently the most conventional method to determine the actual sequence of these important regulators of gene expression. Typical smRNA cloning projects involve the sequencing of hundreds to thousands of smRNA clones that are delimited at their 5' and 3' ends by fixed sequence regions. These primers result from the biochemical protocol used to isolate and convert the smRNA into clonable PCR products. Recently we completed a smRNA cloning project involving tobacco plants, where analysis was required for approximately 700 smRNA sequences. Finding no easily accessible research tool to enter and analyze smRNA sequences we developed Ebbie to assist us with our study. RESULTS: Ebbie is a semi-automated smRNA cloning data processing algorithm, which initially searches for any substring within a DNA sequencing text file, which is flanked by two constant strings. The substring, also termed smRNA or insert, is stored in a MySQL and BlastN database. These inserts are then compared using BlastN to locally installed databases allowing the rapid comparison of the insert to both the growing smRNA database and to other static sequence databases. Our laboratory used Ebbie to analyze scores of DNA sequencing data originating from an smRNA cloning project. Through its built-in instant analysis of all inserts using BlastN, we were able to quickly identify 33 groups of smRNAs from approximately 700 database entries. This clustering allowed the easy identification of novel and highly expressed clusters of smRNAs. Ebbie is available under GNU GPL and currently implemented on http://bioinformatics.org/ebbie/. CONCLUSION: Ebbie was designed for medium sized smRNA cloning projects with about 1,000 database entries. Ebbie can be used for any type of sequence analysis where two constant primer regions flank a sequence of interest. The reliable storage of inserts, and their annotation in a MySQL database, BlastN comparison of new inserts to dynamic and static databases make it a powerful new tool in any laboratory using DNA sequencing. Ebbie also prevents manual mistakes during the excision process and speeds up annotation and data-entry. Once the server is installed locally, its access can be restricted to protect sensitive new DNA sequencing data. Ebbie was primarily designed for smRNA cloning projects, but can be applied to a variety of RNA and DNA cloning projects. H. Alexander Ebhardt, Kay C. Wiese, Peter J. Unrau |
BMC Bioinform. | 2 |
| 2005 | Significance of randomness in P-RnaPredict - a parallel evolutionary algorithm for RNA foldingabstractThis paper presents an extension to P-RnaPredict, a parallel evolutionary algorithm (EA) for RNA folding. The impact of three pseudorandom number generators (PRNGs) on the EA's performance is evaluated. The generators tested included the C standard library PRNG RAND, a parallelized multiplicative congruential generator (MCG), and a parallelized Mersenne Twister (MT). P-RnaPredict was implemented using the message passing interface (MPI) and tested on a 128 node Beowulf cluster. The PRNG comparison testing was performed with four known structures that are 118, 122, 543, and 556 nucleotides in length. PRNGs effects were investigated and predicted structures compared to known structures Kay C. Wiese, Andrew Hendriks, Alain Deschênes, Belgacem Ben Youssef |
Congress on Evolutionary Computation | 1 |
| 2005 | Algorithms for RNA folding: a comparison of dynamic programming and parallel evolutionary algorithmsabstractThis paper presents a comparison of two types of algorithms for RNA secondary structure prediction: an implementation of Nussinov's dynamic programming algorithm (DPA), and P-RnaPredict, a parallel evolutionary algorithm (EA). The research presented here builds on previous work and examines the results from tests of three RNA sequences that are 118, 543, and 784 nucleotides in length. A variety of EA parameter settings were employed based on previous experimentation. Predicted structures were compared to those generated by the Nussinov DPA and to known structures to determine relative accuracy. Results indicate that the EA demonstrated high prediction accuracy and outperformed the Nussinov DPA on all tested sequences Kay C. Wiese, Andrew Hendriks, Jagdeep Poonian |
Congress on Evolutionary Computation | 1 |
| 2005 | The impact of pseudorandom number quality on P-RnaPredict, a parallel genetic algorithm for RNA secondary structure predictionabstractThis paper presents a parallel version of RnaPredict, agenetic algorithm (GA) for RNA secondary structure prediction. The research presented here builds on previous work and examines the impact of three different pseudorandom number generators (PRNGs) on the GA’s performance. The three generators tested are the C standard library PRNG RAND, a parallelized multiplicative congruential generator (MCG), and a parallelized Mersenne Twister (MT). A fully parallel version of RnaPredict using the Message Passing Interface (MPI) was implemented. The PRNG comparison tests were performed with known structures that are 118, 122, 543, and 556 nucleotides in length. The effects of the PRNGs are investigated and the predicted structures are compared to known structures. Kay C. Wiese, Andrew Hendriks, Alain Deschênes, Belgacem Ben Youssef |
GECCO | 1 |
| 2004 | Using stacking-energies (INN and INN-HB) for improving the accuracy of RNA secondary structure prediction with an evolutionary algorithm - a comparison to known structuresabstractThis paper builds on previous research from an EA used to predict secondary structure of RNA molecules. The EA predicts which specific canonical base pairs forms hydrogen bonds and helices. Three new thermodynamic models were integrated into our EA. The first based on a modification to our original base pair model. The last two, INN and INN-HB, add stacking-energies using base pair adjacencies. We have tested RNA sequences of lengths 122, 543, and 1494 nucleotides on a wide variety of operators and parameters settings. The accuracy of the predicted structures is compared to the known structures thus demonstrating the benefits of using stacking-energies in structure prediction. Some other improvements to our EA are also discussed. Alain Deschênes, Kay C. Wiese |
IEEE Congress on Evolutionary Computation | 2 |
| 2004 | Comparison of dynamic programming and evolutionary algorithms for RNA secondary structure predictionabstractThis work builds on previous research from an EA used to predict secondary structure of RNA molecules. The EA has the goal of predicting which canonical base pairs will form hydrogen bonds and helices. The addition of stacking energies, through INN and INN-HB, to our thermodynamic model has enhanced our predictions. We test three RNA sequences of lengths 118, 543, and 784 nucleotides using a variety of previously successful operators and parameter settings. The accuracy of the predicted structures are compared against those generated by the Nussinov DPA and also to known structures. The EA showed high accuracy of prediction especially on short sequences. On all tested sequences, the EA outperforms the Nussinov DPA. Alain Deschênes, Kay C. Wiese, Jagdeep Poonian |
CIBCB | 2 |
| 2004 | A parallel evolutionary algorithm for RNA secondary structure prediction using stacking-energies (INN and INN-HB)abstractThis work presents a coarse-grained distributed genetic algorithm (GA) for RNA secondary structure prediction. This research builds on previous work and contains two new thermodynamic models, INN and INN-HB, which add stacking-energies using base pair adjacencies. Comparison tests were performed against the original serial GA on known structures that are 122, 543, and 784 nucleotides in length on a wide variety of parameter settings. The effects of the new models are investigated, the predicted structures are compared to known structures and the GA is compared against a serial GA with identical models. Both algorithms perform well and are able to predict structures with high accuracy for short sequences. Andrew Hendriks, Alain Deschênes, Kay C. Wiese |
CIBCB | 3 |
| 2003 | A distributed genetic algorithm for RNA secondary structure predictionabstractThis paper presents a new coarse-grained distributed genetic algorithm (GA) for the prediction of the secondary structure of RNA molecules, based largely on a serial permutation-based GA. The benefits of the distributed GA over our existing serial GA are analyzed and demonstrated. We also analyze the impact of the keep-best reproduction (KBR) and roulette wheel selection (STDS) GA replacement techniques. Finally, we verify the increase in convergence speed of our distributed GA. Tests was performed on 241 and 785 nucleotide sequences. Overall, the distributed GA is found to improve upon the serial GA performances, with a much more pronounced impact on the STDS selection strategy. There is also a notable acceleration in convergence speed. Andrew Hendriks, Kay C. Wiese, Edward Glen, Alain Deschênes |
IEEE Congress on Evolutionary Computation | 2 |
| 2003 | Permutation-based RNA secondary structure prediction via a genetic algorithmabstractThis paper presents new results with a permutation-based genetic algorithm (GA) to predict the secondary structure of RNA molecules. More specifically, the proposed algorithm predicts which canonical base pairs forms hydrogen bonds and builds helices, also known as stems. We discuss a GA where a permutation is used to encode the secondary structure of RNA molecules. We have tested RNA sequences of lengths 76, 210, 681, and 785 nucleotides over a wide variety of operators and parameter settings and focus on discussing in depth the results with two crossover operators asymmetric edge recombinations (ASERC) and symmetric edge recombination (SYMERC) that have not been analyzed in this domain previously. We demonstrate that the keep-best reproduction (KBR) operator has similar benefits as in the travelling salesman problem (TSP) domain. We also compare the results of the permutation-based GA with a binary GA, demonstrating the benefits of the newly proposed representation. Kay C. Wiese, Alain Deschênes, Edward Glen |
IEEE Congress on Evolutionary Computation | 1 |
| 2003 | RNA Structure as Permutation: A GA Approach Comparing Different Genetic Sequencing Operators
Kay C. Wiese, Edward Glen |
ISMIS | 1 |
| 2002 | A Permutation Based Genetic Algorithm for RNA Secondary Structure Prediction
Kay C. Wiese, Edward Glen |
HIS | 1 |