Miroslava Cuperlovic-Culf

dblp:57/749 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
7since 2021 · last 2025
0000-0002-9483-8159ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 7 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2025 Computational approaches to enzymatic reaction assignment: a review of methods, validations, and future directions
abstract
Characterizing the proteins and molecules that underpin cellular metabolism is fundamental to advancing our understanding of biological processes. However, the rapidly expanding repertoire of newly identified proteins and metabolites presents significant challenges for experimental characterization and functional analysis. Computational approaches can be used to identify and elucidate catalytic relationships between enzymes and their substrates and provide powerful tools that support biological research and applications in biochemical engineering, and drug discovery. In this review, we describe the problem of reaction assignment for predicting enzymatic reactions leveraging structural, network, and high-throughput experimental data. Also considered are theoretical perspectives motivating the design of computational methods, available resources, and validation techniques. Current and future computational approaches for enzymatic reaction assignment are expected to advance in tandem with technologies for experimental analysis of metabolism, such as metabolomics, and flux-based methods, to expand our understanding of metabolism.
Luke Kennedy, Mary-Ellen Harper, Miroslava Cuperlovic-Culf
Briefings Bioinform.3
2025 A hybrid machine learning framework for functional annotation of mitochondrial glutathione transport and metabolism proteins in cancers
abstract
BACKGROUND: Alterations of metabolism, including changes in mitochondrial metabolism as well as glutathione (GSH) metabolism are a well appreciated hallmark of many cancers. Mitochondrial GSH (mGSH) transport is a poorly characterized aspect of GSH metabolism, which we investigate in the context of cancer. Existing functional annotation approaches from machine (ML) or deep learning (DL) models based only on protein sequences, were unable to annotate functions in biological contexts. RESULTS: We develop a flexible ML framework for functional annotation from diverse feature data. This hybrid ML framework leverages cancer cell line multi-omics data and other biological knowledge data as features, to uncover potential genes involved in mGSH metabolism and membrane transport in cancers. This framework achieves strong performance across functional annotation tasks and several cell line and primary tumor cancer samples. For our application, classification models predict the known mGSH transporter SLC25A39 but not SLC25A40 as being highly probably related to mGSH metabolism in cancers. SLC25A10, SLC25A50, and orphan SLC25A24, SLC25A43 are predicted to be associated with mGSH metabolism in multiple biological contexts and structural analysis of these proteins reveal similarities in potential substrate binding regions to the binding residues of SLC25A39. CONCLUSION: These findings have implications for a better understanding of cancer cell metabolism and novel therapeutic targets with respect to GSH metabolism through potential novel functional annotations of genes. The hybrid ML framework proposed here can be applied to other biological function classifications or multi-omics datasets to generate hypotheses in various biological contexts. Code and a tutorial for generating models and predictions in this framework are available at: https://github.com/lkenn012/mGSH_cancerClassifiers .
Luke Kennedy, Jagdeep K. Sandhu, Mary-Ellen Harper, Miroslava Cuperlovic-Culf
BMC Bioinform.4
2023 Signed Distance Correlation (SiDCo): an online implementation of distance correlation and partial distance correlation for data-driven network analysis
abstract
MOTIVATION: There is a need for easily accessible implementations that measure the strength of both linear and non-linear relationships between metabolites in biological systems as an approach for data-driven network development. While multiple tools implement linear Pearson and Spearman methods, there are no such tools that assess distance correlation. RESULTS: We present here SIgned Distance COrrelation (SiDCo). SiDCo is a GUI platform for calculation of distance correlation in omics data, measuring linear and non-linear dependencies between variables, as well as correlation between vectors of different lengths, e.g. different sample sizes. By combining the sign of the overall trend from Pearson's correlation with distance correlation values, we further provide a novel "signed distance correlation" of particular use in metabolomic and lipidomic analyses. Distance correlations can be selected as one-to-one or one-to-all correlations, showing relationships between each feature and all other features one at a time or in combination. Additionally, we implement "partial distance correlation," calculated using the Gaussian Graphical model approach adapted to distance covariance. Our platform provides an easy-to-use software implementation that can be applied to the investigation of any dataset. AVAILABILITY AND IMPLEMENTATION: The SiDCo software application is freely available at https://complimet.ca/sidco. Supplementary help pages are provided at https://complimet.ca/sidco. Supplementary Material shows an example of an application of SiDCo in metabolomics.
Francesco Monti, David Stewart, Anuradha Surendra, Irina Alecu, Thao Nguyen-Tran, Steffany A. L. Bennett, Miroslava Cuperlovic-Culf
Bioinform.7
2023 DAPTEV: Deep aptamer evolutionary modelling for COVID-19 drug design
abstract
Typical drug discovery and development processes are costly, time consuming and often biased by expert opinion. Aptamers are short, single-stranded oligonucleotides (RNA/DNA) that bind to target proteins and other types of biomolecules. Compared with small-molecule drugs, aptamers can bind to their targets with high affinity (binding strength) and specificity (uniquely interacting with the target only). The conventional development process for aptamers utilizes a manual process known as Systematic Evolution of Ligands by Exponential Enrichment (SELEX), which is costly, slow, dependent on library choice and often produces aptamers that are not optimized. To address these challenges, in this research, we create an intelligent approach, named DAPTEV, for generating and evolving aptamer sequences to support aptamer-based drug discovery and development. Using the COVID-19 spike protein as a target, our computational results suggest that DAPTEV is able to produce structurally complex aptamers with strong binding affinities.
Cameron Andress, Kalli Kappel, Marcus Elbert Villena, Miroslava Cuperlovic-Culf, Hongbin Yan
PLoS Comput. Biol.4
2022 BATL: Bayesian annotations for targeted lipidomics
abstract
MOTIVATION: Bioinformatic tools capable of annotating, rapidly and reproducibly, large, targeted lipidomic datasets are limited. Specifically, few programs enable high-throughput peak assessment of liquid chromatography-electrospray ionization tandem mass spectrometry data acquired in either selected or multiple reaction monitoring modes. RESULTS: We present here Bayesian Annotations for Targeted Lipidomics, a Gaussian naïve Bayes classifier for targeted lipidomics that annotates peak identities according to eight features related to retention time, intensity, and peak shape. Lipid identification is achieved by modeling distributions of these eight input features across biological conditions and maximizing the joint posterior probabilities of all peak identities at a given transition. When applied to sphingolipid and glycerophosphocholine selected reaction monitoring datasets, we demonstrate over 95% of all peaks are rapidly and correctly identified. AVAILABILITY AND IMPLEMENTATION: BATL software is freely accessible online at https://complimet.ca/batl/ and is compatible with Safari, Firefox, Chrome and Edge. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Justin G. Chitpin, Anuradha Surendra, Thao T. Nguyen, Graeme Taylor, Irina Alecu, Roberto Ortega, Julianna J. Tomlinson, Angela M. Crawley, Michaeline McGuinty, Michael G. Schlossmacher, Rachel Saunders-Pullman, Miroslava Cuperlovic-Culf, Steffany A. L. Bennett, Theodore J. Perkins
Bioinform.13
2022 METAbolomics data Balancing with Over-sampling Algorithms (META-BOA): an online resource for addressing class imbalance
abstract
MOTIVATION: Class imbalance, or unequal sample sizes between classes, is an increasing concern in machine learning for metabolomic and lipidomic data mining, which can result in overfitting for the over-represented class. Numerous methods have been developed for handling class imbalance, but they are not readily accessible to users with limited computational experience. Moreover, there is no resource that enables users to easily evaluate the effect of different over-sampling algorithms. RESULTS: METAbolomics data Balancing with Over-sampling Algorithms (META-BOA) is a web-based application that enables users to select between four different methods for class balancing, followed by data visualization and classification of the sample to observe the augmentation effects. META-BOA outputs a newly balanced dataset, generating additional samples in the minority class, according to the user's choice of Synthetic Minority Over-sampling Technique (SMOTE), Borderline-SMOTE (BSMOTE), Adaptive Synthetic (ADASYN) or Random Over-Sampling Examples (ROSE). To present the effect of over-sampling on the data META-BOA further displays both principal component analysis and t-distributed stochastic neighbor embedding visualization of data pre- and post-over-sampling. Random forest classification is utilized to compare sample classification in both the original and balanced datasets, enabling users to select the most appropriate method for their further analyses. AVAILABILITY AND IMPLEMENTATION: META-BOA is available at https://complimet.ca/meta-boa. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Emily Hashimoto-Roth, Anuradha Surendra, Mathieu Lavallée-Adam, Steffany A. L. Bennett, Miroslava Cuperlovic-Culf
Bioinform.5
2021 SMILE: systems metabolomics using interpretable learning and evolution
abstract
BACKGROUND: Direct link between metabolism and cell and organism phenotype in health and disease makes metabolomics, a high throughput study of small molecular metabolites, an essential methodology for understanding and diagnosing disease development and progression. Machine learning methods have seen increasing adoptions in metabolomics thanks to their powerful prediction abilities. However, the "black-box" nature of many machine learning models remains a major challenge for wide acceptance and utility as it makes the interpretation of decision process difficult. This challenge is particularly predominant in biomedical research where understanding of the underlying decision making mechanism is essential for insuring safety and gaining new knowledge. RESULTS: In this article, we proposed a novel computational framework, Systems Metabolomics using Interpretable Learning and Evolution (SMILE), for supervised metabolomics data analysis. Our methodology uses an evolutionary algorithm to learn interpretable predictive models and to identify the most influential metabolites and their interactions in association with disease. Moreover, we have developed a web application with a graphical user interface that can be used for easy analysis, interpretation and visualization of the results. Performance of the method and utilization of the web interface is shown using metabolomics data for Alzheimer's disease. CONCLUSIONS: SMILE was able to identify several influential metabolites on AD and to provide interpretable predictive models that can be further used for a better understanding of the metabolic background of AD. SMILE addresses the emerging issue of interpretability and explainability in machine learning, and contributes to more transparent and powerful applications of machine learning in bioinformatics.
Chengyuan Sha, Miroslava Cuperlovic-Culf
BMC Bioinform.2
2019 PROAFTN Classifier for Feature Selection with Application to Alzheimer Metabolomics Data Analysis
abstract
Early and accurate Alzheimer’s disease (AD) diagnosis remains a challenge. Recently, increasing efforts have been focused towards utilization of metabolomics data for the discovery of biomarkers for screening and diagnosis of AD. Several machine learning approaches were explored for classifying the blood metabolomics profiles of cognitively healthy and AD patients. Differentiation between AD, mild cognitive impairment (MCI) and cognitively healthy subjects remains difficult. In this paper, we propose a new machine learning approach for the selection of a subset of features that provide an improvement in classification rates between these three levels of cognitive disorders. Our experimental results demonstrate that utilization of these selected metabolic markers improves the performance of several classifiers in comparison to the classification accuracy obtained for the complete metabolomics dataset. The obtained results indicate that our algorithms are effective in discovering a panel of biomarkers of AD and MCI from metabolomics data suggesting the possibility to develop a noninvasive blood diagnostic technique of AD and MCI.
Nabil Belacel, Miroslava Cuperlovic-Culf
Int. J. Pattern Recognit. Artif. Intell.2
2011 MetaboHunter: an automatic approach for identification of metabolites from 1H-NMR spectra of complex mixtures
abstract
BACKGROUND: One-dimensional 1H-NMR spectroscopy is widely used for high-throughput characterization of metabolites in complex biological mixtures. However, the accurate identification of individual compounds is still a challenging task, particularly in spectral regions with higher peak densities. The need for automatic tools to facilitate and further improve the accuracy of such tasks, while using increasingly larger reference spectral libraries becomes a priority of current metabolomics research. RESULTS: We introduce a web server application, called MetaboHunter, which can be used for automatic assignment of 1H-NMR spectra of metabolites. MetaboHunter provides methods for automatic metabolite identification based on spectra or peak lists with three different search methods and with possibility for peak drift in a user defined spectral range. The assignment is performed using as reference libraries manually curated data from two major publicly available databases of NMR metabolite standard measurements (HMDB and MMCD). Tests using a variety of synthetic and experimental spectra of single and multi metabolite mixtures show that MetaboHunter is able to identify, in average, more than 80% of detectable metabolites from spectra of synthetic mixtures and more than 50% from spectra corresponding to experimental mixtures. This work also suggests that better scoring functions improve by more than 30% the performance of MetaboHunter's metabolite identification methods. CONCLUSIONS: MetaboHunter is a freely accessible, easy to use and user friendly 1H-NMR-based web server application that provides efficient data input and pre-processing, flexible parameter settings, fast and automatic metabolite fingerprinting and results visualization via intuitive plotting and compound peak hit maps. Compared to other published and freely accessible metabolomics tools, MetaboHunter implements three efficient methods to search for metabolites in manually curated data from two reference libraries.
Dan C. Tulpan, Serge Léger, Luc Belliveau, Adrian Culf, Miroslava Cuperlovic-Culf
BMC Bioinform.5
2004 Fuzzy J-Means and VNS methods for clustering genes from microarray data
abstract
MOTIVATION: In the interpretation of gene expression data from a group of microarray experiments that include samples from either different patients or conditions, special consideration must be given to the pleiotropic and epistatic roles of genes, as observed in the variation of gene coexpression patterns. Crisp clustering methods assign each gene to one cluster, thereby omitting information about the multiple roles of genes. RESULTS: Here, we present the application of a local search heuristic, Fuzzy J-Means, embedded into the variable neighborhood search metaheuristic for the clustering of microarray gene expression data. We show that for all the datasets studied this algorithm outperforms the standard Fuzzy C-Means heuristic. Different methods for the utilization of cluster membership information in determining gene coregulation are presented. The clustering and data analyses were performed on simulated datasets as well as experimental cDNA microarray data for breast cancer and human blood from the Stanford Microarray Database. AVAILABILITY: The source code of the clustering software (C programming language) is freely available from [email protected]
Nabil Belacel, Miroslava Cuperlovic-Culf, Mark Laflamme, Rodney Ouellette
Bioinform.2