Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jonathan Raad

dblp:255/9148 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
4since 2021 · last 2024
0000-0002-9585-1531ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
4 papers
Bioinformatics and computational biology · 97% Medical and health informatics · 3%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › transcriptomics
non-coding RNA analysis
1.022022
miRe2e: a full end-to-end deep model based on transformers for prediction of pre-miRNAs · Bioinform. 2022
Complexity measures of the mature miRNA for improving pre-miRNAs prediction · Bioinform. 2020
Bioinformatics and computational biology › transcriptomics › non-coding RNA analysis
pre-miRNA prediction
1.022022
miRe2e: a full end-to-end deep model based on transformers for prediction of pre-miRNAs · Bioinform. 2022
Complexity measures of the mature miRNA for improving pre-miRNAs prediction · Bioinform. 2020
Bioinformatics and computational biology › genomics
viral genomics
0.512021
Novel SARS-CoV-2 encoded small RNAs in the passage to humans · Bioinform. 2021
Bioinformatics and computational biology
biomedical text mining
0.412020
DL4papers: a deep learning approach for the automatic interpretation of scientific articles · Bioinform. 2020
Bioinformatics and computational biology › biomedical text mining
relation extraction
0.412020
DL4papers: a deep learning approach for the automatic interpretation of scientific articles · Bioinform. 2020
Information retrieval › retrieval models
ranked retrieval
0.412020
DL4papers: a deep learning approach for the automatic interpretation of scientific articles · Bioinform. 2020
Bioinformatics and computational biology › genomics › machine learning for genomics
deep learning for genomics
0.212022
miRe2e: a full end-to-end deep model based on transformers for prediction of pre-miRNAs · Bioinform. 2022
Bioinformatics and computational biology
transcriptomics
0.112021
Novel SARS-CoV-2 encoded small RNAs in the passage to humans · Bioinform. 2021
Medical and health informatics
precision medicine
0.112020
DL4papers: a deep learning approach for the automatic interpretation of scientific articles · Bioinform. 2020

Methods — techniques the papers use, named apart from their topics

deep learning · 1.8text mining · 0.9transformer · 0.6attention mechanism · 0.6small RNA-seq · 0.5supervised machine learning · 0.4
YearPublicationVenuePosition
2024 sincFold: end-to-end learning of short- and long-range interactions in RNA secondary structure
abstract
MOTIVATION: Coding and noncoding RNA molecules participate in many important biological processes. Noncoding RNAs fold into well-defined secondary structures to exert their functions. However, the computational prediction of the secondary structure from a raw RNA sequence is a long-standing unsolved problem, which after decades of almost unchanged performance has now re-emerged due to deep learning. Traditional RNA secondary structure prediction algorithms have been mostly based on thermodynamic models and dynamic programming for free energy minimization. More recently deep learning methods have shown competitive performance compared with the classical ones, but there is still a wide margin for improvement. RESULTS: In this work we present sincFold, an end-to-end deep learning approach, that predicts the nucleotides contact matrix using only the RNA sequence as input. The model is based on 1D and 2D residual neural networks that can learn short- and long-range interaction patterns. We show that structures can be accurately predicted with minimal physical assumptions. Extensive experiments were conducted on several benchmark datasets, considering sequence homology and cross-family validation. sincFold was compared with classical methods and recent deep learning models, showing that it can outperform the state-of-the-art methods.
Leandro A. Bugnon, Leandro E. Di Persia, Matias Gerard, Jonathan Raad, Santiago Prochetto, Emilio Fenoy, Uciel Chorostecki, Federico Ariel, Georgina Stegmayer, Diego H. Milone
Briefings Bioinform.4
2022 Secondary structure prediction of long noncoding RNA: review and experimental comparison of existing approaches
abstract
MOTIVATION: In contrast to messenger RNAs, the function of the wide range of existing long noncoding RNAs (lncRNAs) largely depends on their structure, which determines interactions with partner molecules. Thus, the determination or prediction of the secondary structure of lncRNAs is critical to uncover their function. Classical approaches for predicting RNA secondary structure have been based on dynamic programming and thermodynamic calculations. In the last 4 years, a growing number of machine learning (ML)-based models, including deep learning (DL), have achieved breakthrough performance in structure prediction of biomolecules such as proteins and have outperformed classical methods in short transcripts folding. Nevertheless, the accurate prediction for lncRNA still remains far from being effectively solved. Notably, the myriad of new proposals has not been systematically and experimentally evaluated. RESULTS: In this work, we compare the performance of the classical methods as well as the most recently proposed approaches for secondary structure prediction of RNA sequences using a unified and consistent experimental setup. We use the publicly available structural profiles for 3023 yeast RNA sequences, and a novel benchmark of well-characterized lncRNA structures from different species. Moreover, we propose a novel metric to assess the predictive performance of methods, exclusively based on the chemical probing data commonly used for profiling RNA structures, avoiding any potential bias incorporated by computational predictions when using dot-bracket references. Our results provide a comprehensive comparative assessment of existing methodologies, and a novel and public benchmark resource to aid in the development and comparison of future approaches. AVAILABILITY: Full source code and benchmark datasets are available at: https://github.com/sinc-lab/lncRNA-folding. CONTACT: [email protected].
Leandro A. Bugnon, Alejandro Edera, Santiago Prochetto, Matias Gerard, Jonathan Raad, Emilio Fenoy, Mariano Rubiolo, Uciel Chorostecki, Toni Gabaldón, Federico Ariel, Leandro E. Di Persia, Diego H. Milone, Georgina Stegmayer
Briefings Bioinform.5
2022 miRe2e: a full end-to-end deep model based on transformers for prediction of pre-miRNAs
abstract
MOTIVATION: MicroRNAs (miRNAs) are small RNA sequences with key roles in the regulation of gene expression at post-transcriptional level in different species. Accurate prediction of novel miRNAs is needed due to their importance in many biological processes and their associations with complicated diseases in humans. Many machine learning approaches were proposed in the last decade for this purpose, but requiring handcrafted features extraction to identify possible de novo miRNAs. More recently, the emergence of deep learning (DL) has allowed the automatic feature extraction, learning relevant representations by themselves. However, the state-of-art deep models require complex pre-processing of the input sequences and prediction of their secondary structure to reach an acceptable performance. RESULTS: In this work, we present miRe2e, the first full end-to-end DL model for pre-miRNA prediction. This model is based on Transformers, a neural architecture that uses attention mechanisms to infer global dependencies between inputs and outputs. It is capable of receiving the raw genome-wide data as input, without any pre-processing nor feature engineering. After a training stage with known pre-miRNAs, hairpin and non-harpin sequences, it can identify all the pre-miRNA sequences within a genome. The model has been validated through several experimental setups using the human genome, and it was compared with state-of-the-art algorithms obtaining 10 times better performance. AVAILABILITY AND IMPLEMENTATION: Webdemo available at https://sinc.unl.edu.ar/web-demo/miRe2e/ and source code available for download at https://github.com/sinc-lab/miRe2e. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jonathan Raad, Leandro A. Bugnon, Diego H. Milone, Georgina Stegmayer
Bioinform.1
2021 Novel SARS-CoV-2 encoded small RNAs in the passage to humans
abstract
MOTIVATION: The Severe Acute Respiratory Syndrome-Coronavirus 2 (SARS-CoV-2) has recently emerged as the responsible for the pandemic outbreak of the coronavirus disease 2019. This virus is closely related to coronaviruses infecting bats and Malayan pangolins, species suspected to be an intermediate host in the passage to humans. Several genomic mutations affecting viral proteins have been identified, contributing to the understanding of the recent animal-to-human transmission. However, the capacity of SARS-CoV-2 to encode functional putative microRNAs (miRNAs) remains largely unexplored. RESULTS: We have used deep learning to discover 12 candidate stem-loop structures hidden in the viral protein-coding genome. Among the precursors, the expression of eight mature miRNAs-like sequences was confirmed in small RNA-seq data from SARS-CoV-2 infected human cells. Predicted miRNAs are likely to target a subset of human genes of which 109 are transcriptionally deregulated upon infection. Remarkably, 28 of those genes potentially targeted by SARS-CoV-2 miRNAs are down-regulated in infected human cells. Interestingly, most of them have been related to respiratory diseases and viral infection, including several afflictions previously associated with SARS-CoV-1 and SARS-CoV-2. The comparison of SARS-CoV-2 pre-miRNA sequences with those from bat and pangolin coronaviruses suggests that single nucleotide mutations could have helped its progenitors jumping inter-species boundaries, allowing the gain of novel mature miRNAs targeting human mRNAs. Our results suggest that the recent acquisition of novel miRNAs-like sequences in the SARS-CoV-2 genome may have contributed to modulate the transcriptional reprograming of the new host upon infection. AVAILABILITY AND IMPLEMENTATION: https://github.com/sinc-lab/sarscov2-mirna-discovery. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Gabriela Alejandra Merino, Jonathan Raad, Leandro A. Bugnon, Cristian A. Yones, Laura Kamenetzky, Juan Claus, Federico Ariel, Diego H. Milone, Georgina Stegmayer
Bioinform.2
2020 DL4papers: a deep learning approach for the automatic interpretation of scientific articles
abstract
MOTIVATION: In precision medicine, next-generation sequencing and novel preclinical reports have led to an increasingly large amount of results, published in the scientific literature. However, identifying novel treatments or predicting a drug response in, for example, cancer patients, from the huge amount of papers available remains a laborious and challenging work. This task can be considered a text mining problem that requires reading a lot of academic documents for identifying a small set of papers describing specific relations between key terms. Due to the infeasibility of the manual curation of these relations, computational methods that can automatically identify them from the available literature are urgently needed. RESULTS: We present DL4papers, a new method based on deep learning that is capable of analyzing and interpreting papers in order to automatically extract relevant relations between specific keywords. DL4papers receives as input a query with the desired keywords, and it returns a ranked list of papers that contain meaningful associations between the keywords. The comparison against related methods showed that our proposal outperformed them in a cancer corpus. The reliability of the DL4papers output list was also measured, revealing that 100% of the first two documents retrieved for a particular search have relevant relations, in average. This shows that our model can guarantee that in the top-2 papers of the ranked list, the relation can be effectively found. Furthermore, the model is capable of highlighting, within each document, the specific fragments that have the associations of the input keywords. This can be very useful in order to pay attention only to the highlighted text, instead of reading the full paper. We believe that our proposal could be used as an accurate tool for rapidly identifying relationships between genes and their mutations, drug responses and treatments in the context of a certain disease. This new approach can certainly be a very useful and valuable resource for the advancement of the precision medicine field. AVAILABILITY AND IMPLEMENTATION: A web-demo is available at: http://sinc.unl.edu.ar/web-demo/dl4papers/. Full source code and data are available at: https://sourceforge.net/projects/sourcesinc/files/dl4papers/. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Leandro A. Bugnon, Cristian A. Yones, Jonathan Raad, Matias Gerard, Mariano Rubiolo, Gabriela Alejandra Merino, Milton Pividori, Leandro E. Di Persia, Diego H. Milone, Georgina Stegmayer
Bioinform.3
2020 Complexity measures of the mature miRNA for improving pre-miRNAs prediction
abstract
MOTIVATION: The discovery of microRNA (miRNA) in the last decade has certainly changed the understanding of gene regulation in the cell. Although a large number of algorithms with different features have been proposed, they still predict an impractical amount of false positives. Most of the proposed features are based on the structure of precursors of the miRNA only, not considering the important and relevant information contained in the mature miRNA. Such new kind of features could certainly improve the performance of the predictors of new miRNAs. RESULTS: This paper presents three new features that are based on the sequence information contained in the mature miRNA. We will show how these new features, when used by a classical supervised machine learning approach as well as by more recent proposals based on deep learning, improve the prediction performance in a significant way. Moreover, several experimental conditions were defined and tested to evaluate the novel features impact in situations close to genome-wide analysis. The results show that the incorporation of new features based on the mature miRNA allows to improve the detection of new miRNAs independently of the classifier used. AVAILABILITY AND IMPLEMENTATION: https://sourceforge.net/projects/sourcesinc/files/cplxmirna/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jonathan Raad, Georgina Stegmayer, Diego H. Milone
Bioinform.1
2019 Predicting novel microRNA: a comprehensive comparison of machine learning approaches
abstract
MOTIVATION: The importance of microRNAs (miRNAs) is widely recognized in the community nowadays because these short segments of RNA can play several roles in almost all biological processes. The computational prediction of novel miRNAs involves training a classifier for identifying sequences having the highest chance of being precursors of miRNAs (pre-miRNAs). The big issue with this task is that well-known pre-miRNAs are usually few in comparison with the hundreds of thousands of candidate sequences in a genome, which results in high class imbalance. This imbalance has a strong influence on most standard classifiers, and if not properly addressed in the model and the experiments, not only performance reported can be completely unrealistic but also the classifier will not be able to work properly for pre-miRNA prediction. Besides, another important issue is that for most of the machine learning (ML) approaches already used (supervised methods), it is necessary to have both positive and negative examples. The selection of positive examples is straightforward (well-known pre-miRNAs). However, it is difficult to build a representative set of negative examples because they should be sequences with hairpin structure that do not contain a pre-miRNA. RESULTS: This review provides a comprehensive study and comparative assessment of methods from these two ML approaches for dealing with the prediction of novel pre-miRNAs: supervised and unsupervised training. We present and analyze the ML proposals that have appeared during the past 10 years in literature. They have been compared in several prediction tasks involving two model genomes and increasing imbalance levels. This work provides a review of existing ML approaches for pre-miRNA prediction and fair comparisons of the classifiers with same features and data sets, instead of just a revision of published software tools. The results and the discussion can help the community to select the most adequate bioinformatics approach according to the prediction task at hand. The comparative results obtained suggest that from low to mid-imbalance levels between classes, supervised methods can be the best. However, at very high imbalance levels, closer to real case scenarios, models including unsupervised and deep learning can provide better performance.
Georgina Stegmayer, Leandro E. Di Persia, Mariano Rubiolo, Matias Gerard, Milton Pividori, Cristian A. Yones, Leandro A. Bugnon, Tadeo Rodriguez, Jonathan Raad, Diego H. Milone
Briefings Bioinform.9