EDBT 2026 Demo / reviewers in the wild / expert
Alexandre Rossi Paschoal
dblp:223/7311
· DBLP profile ↗
11ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0002-8887-0582ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | mirtronDB 2.0: enhanced database with novel mirtron discoveriesabstractMOTIVATION: MirtronDB provides a comprehensive and up-to-date resource for advancing mirtron research within RNA biology. Therefore, maintaining a specialized and continuously updated resource for mirtrons is essential to support ongoing discoveries and to serve as a key reference for researchers investigating the roles of mirtrons. RESULTS: Here, we present mirtronDB 2.0, an enhanced version that expands both content and functionality. This version integrates mirtron data published between 2017 and 2025, increasing the number of documented mirtrons across various species. In addition, it incorporates newly predicted mirtrons identified through a robust pipeline that combines advanced bioinformatics and machine learning approaches, with specific coverage of six mammalian species. We have introduced new website features, including an interactive dashboard to enhance usability and facilitate intuitive data exploration. These rigorous updates consolidate mirtronDB as a key resource for mirtron to the RNA biology community. AVAILABILITY AND IMPLEMENTATION: mirtronDB can be found under http://mirtrondb.cp.utfpr.edu.br/. The complete content of Database 2.0 and the source code for the analyses are also freely available in the FigShare repository: https://figshare.com/articles/dataset/MirtronDB_version2/29344775. Fabiana Rodrigues de Góes, Matheus Fujimura Soares, Vitor Gregorio, Bruno Thiago de Lima Nichio, Alisson Gaspar Chiquitto, Flavia Lombardi Lopes, Mark Basham, Douglas Silva Domingues, Alexandre Rossi Paschoal |
Bioinform. | 9 |
| 2025 | DeepSEA: an alignment-free explainable approach to annotate antimicrobial resistance proteinsabstractAntimicrobial resistance (AMR) is one of the most concerning modern threats as it places a greater burden on health systems than HIV and malaria combined. Current surveillance strategies for tracking antimicrobial resistance (AMR) rely on genomic comparisons and depend on sequence alignment with strict similarity cutoffs of greater than 95%. Therefore, these methods have high false-negative error rates due to a lack of reference sequences with a representative coverage of AMR protein diversity. Deep learning has been used as an alternative to sequence alignment, as artificial neural networks can extract abstract features from data, thereby limiting the need for sequence comparisons. Here, a convolutional neural network (CNN) was trained to differentiate between antimicrobial resistance proteins and non-resistance proteins, and to annotate them in nine resistance classes. Our model demonstrated higher recall values (> 0.9) than the alignment-based approach for all protein classes tested. Additionally, our CNN architecture allowed us to investigate internal states and explain the model classification regarding protein domain feature importance related to antimicrobial molecule inactivation. Finally, we built an open-source bioinformatic tool ( https://github.com/computational-chemical-biology/DeepSEA-project ) that can be used to annotate antimicrobial resistance proteins and provide information on protein domains without sequence alignment. Tiago Cabral Borelli, Alexandre Rossi Paschoal, Ricardo Roberto da Silva |
BMC Bioinform. | 2 |
| 2024 | Classification of LTR Retrotransposons via Interaction PredictionabstractTransposable Elements (TEs) are genetic sequences that can relocate within the genome, promoting genetic diversity. In eukaryotes, TEs are classified into classes, subclasses, orders, superfamilies, families, and subfamilies. LTR retrotransposons (LTR-RT) constitute an order in this taxonomy. The main objective of this study is to investigate the classification of LTR retrotransposons at the superfamily level. Predictive Bi-Clustering Trees (PBCTs) were used to predict interactions between LTR-RT sequences and conserved protein domains to achieve this. Two datasets were used to investigate the relationships among different superfamilies. The first dataset contained LTR retrotransposon sequences assigned to Copia, Gypsy, and Bel-Pao superfamilies, while the second dataset included consensus sequences of the conserved domains for each superfamily. Thus, the PBCT decision tree tests could relate to the LTR-RT sequence and conserved domain attributes. In the classification process, interaction is interpreted as either the presence or absence of a domain in a given LTR-RT sequence. The sequence is then classified into the superfamily with the most predicted domains. Precision-recall curves were adopted as evaluation metrics for the method, and its performance was compared to some of the most commonly used models in the task of transposable element classification. The experiments conducted on D. melanogaster and A. thaliana showed that PBCTs are promising and comparable to other methods, especially in classifying the Gypsy superfamily. Silvana C. S. Cardoso, Douglas Silva Domingues, Alexandre Rossi Paschoal, Carlos Fischer, Ricardo Cerri |
CIBCB | 3 |
| 2023 | Inpactor2: a software based on deep learning to identify and classify LTR-retrotransposons in plant genomesabstractLTR-retrotransposons are the most abundant repeat sequences in plant genomes and play an important role in evolution and biodiversity. Their characterization is of great importance to understand their dynamics. However, the identification and classification of these elements remains a challenge today. Moreover, current software can be relatively slow (from hours to days), sometimes involve a lot of manual work and do not reach satisfactory levels in terms of precision and sensitivity. Here we present Inpactor2, an accurate and fast application that creates LTR-retrotransposon reference libraries in a very short time. Inpactor2 takes an assembled genome as input and follows a hybrid approach (deep learning and structure-based) to detect elements, filter partial sequences and finally classify intact sequences into superfamilies and, as very few tools do, into lineages. This tool takes advantage of multi-core and GPU architectures to decrease execution times. Using the rice genome, Inpactor2 showed a run time of 5 minutes (faster than other tools) and has the best accuracy and F1-Score of the tools tested here, also having the second best accuracy and specificity only surpassed by EDTA, but achieving 28% higher sensitivity. For large genomes, Inpactor2 is up to seven times faster than other available bioinformatics tools. Simon Orozco-Arias, Luis Humberto López-Murillo, Mariana S. Candamil-Cortes, Maradey Arias, Paula A. Jaimes, Alexandre Rossi Paschoal, Reinel Tabares-Soto, Gustavo A. Isaza, Romain Guyot |
Briefings Bioinform. | 6 |
| 2022 | Impact of sequencing technologies on long non-coding RNA computational identificationabstractThe correct annotation of non-coding RNAs, especially long non-coding RNAs (lncRNAs), is still a critial challenge in genome analyses due to their highly heterogeneous characteristics. Due to this heterogeneity, transcriptome data sources can be an important factor that might affect lncRNA annotation quality. Long-read technologies now bring the potential to improve the quality of transcriptome annotation, specially in genome entities that are not “classic” coding genes. However, there is a gap regarding benchmarking studies that test if the direct use of lncRNA predictors in long-reads makes more precise identification of these transcripts. Considering that lncRNA identification tools were not trained with these reads, our study want to address: how is the performance of these tools? Are they also able to efficiently identify lncRNAs? For this, we used short and long-read data from human and selected plants transcriptomes to test our questions. We can provide evidence of where and how to make potential better approaches for the lncRNA annotation by understanding these issues. Alisson Gaspar Chiquitto, Lucas Otávio L. Silva, Liliane S. Oliveira, Douglas Silva Domingues, Alexandre Rossi Paschoal |
BIBM | 5 |
| 2021 | Feature extraction approaches for biological sequences: a comparative study of mathematical featuresabstractAs consequence of the various genomic sequencing projects, an increasing volume of biological sequence data is being produced. Although machine learning algorithms have been successfully applied to a large number of genomic sequence-related problems, the results are largely affected by the type and number of features extracted. This effect has motivated new algorithms and pipeline proposals, mainly involving feature extraction problems, in which extracting significant discriminatory information from a biological set is challenging. Considering this, our work proposes a new study of feature extraction approaches based on mathematical features (numerical mapping with Fourier, entropy and complex networks). As a case study, we analyze long non-coding RNA sequences. Moreover, we separated this work into three studies. First, we assessed our proposal with the most addressed problem in our review, e.g. lncRNA and mRNA; second, we also validate the mathematical features in different classification problems, to predict the class of lncRNA, e.g. circular RNAs sequences; third, we analyze its robustness in scenarios with imbalanced data. The experimental results demonstrated three main contributions: first, an in-depth study of several mathematical features; second, a new feature extraction pipeline; and third, its high performance and robustness for distinct RNA sequence classification. Availability:https://github.com/Bonidia/FeatureExtraction_BiologicalSequences. Robson Bonidia, Lucas Dias H. Sampaio, Douglas Silva Domingues, Alexandre Rossi Paschoal, Fabrício Martins Lopes, André C. P. L. F. de Carvalho, Danilo Sipoli Sanches |
Briefings Bioinform. | 4 |
| 2021 | TERL: classification of transposable elements by convolutional neural networksabstractTransposable elements (TEs) are the most represented sequences occurring in eukaryotic genomes. Few methods provide the classification of these sequences into deeper levels, such as superfamily level, which could provide useful and detailed information about these sequences. Most methods that classify TE sequences use handcrafted features such as k-mers and homology-based search, which could be inefficient for classifying non-homologous sequences. Here we propose an approach, called transposable elements pepresentation learner (TERL), that preprocesses and transforms one-dimensional sequences into two-dimensional space data (i.e., image-like data of the sequences) and apply it to deep convolutional neural networks. This classification method tries to learn the best representation of the input data to classify it correctly. We have conducted six experiments to test the performance of TERL against other methods. Our approach obtained macro mean accuracies and F1-score of 96.4% and 85.8% for superfamilies and 95.7% and 91.5% for the order sequences from RepBase, respectively. We have also obtained macro mean accuracies and F1-score of 95.0% and 70.6% for sequences from seven databases into superfamily level and 89.3% and 73.9% for the order level, respectively. We surpassed accuracy, recall and specificity obtained by other methods on the experiment with the classification of order level sequences from seven databases and surpassed by far the time elapsed of any other method for all experiments. Therefore, TERL can learn how to predict any hierarchical level of the TEs classification system and is about 20 times and three orders of magnitude faster than TEclass and PASTEC, respectively https://github.com/muriloHoracio/TERL. Contact:[email protected]. Murilo Horacio Pereira da Cruz, Douglas Silva Domingues, Priscila T. M. Saito, Alexandre Rossi Paschoal, Pedro Henrique Bugatti |
Briefings Bioinform. | 4 |
| 2020 | Comparison tools for lncRNA identification: analysis among plants and humansabstractThis article has as its main objective the evaluation of the differences between long non-coding RNAs of plants and humans. Long non-coding RNAs are also known as lncRNAs. The lncRNAS belong to the class of RNAs that do not encode proteins and are related to several biological functions, such as chromatin modifications, post-transcriptional regulation and mainly in the different development processes of diseases such as cancer. In this work, we want to verify the existence of differences in lncRNAs in plants and humans using state-of-the-art approaches to identify lncRNAs. The main reason for the study is that there are differences between the miNAs (small ncRNAs) of plants and humans, whether in biological or computational characteristics, for lncRNAs it is still an open question. To answer this question, this paper proposes to show the results of two ncRNAS prediction tools, trained with humans, and which are widely used for lncRNA prediction: CPC2 and CPAT. We will also show results from tools used to predict lncRNAS in plants, which are trained with plant data: the RNAplonc, the PlncPRO tool that contains two versions, one for monocot and one for dicot and the LGC tool that was trained with plants and humans. The results of tools trained with human data will also be displayed: PLEK, CPPRED and PredLnc-GFStack. These eight tools were applied in two sets of tests, one composed of eight species of plants (Amborella trichopoda, Brachypodium distachyon, Citrus sinensis, Manihot esculenta, Ricinus communis, Solanum tuberosum, Sorghum bicolor, Zea mays) and the other composed of human lncRNAS. Tatianne da Costa Negri, Alexandre Rossi Paschoal, Wonder Alexandre Luz Alves |
CIBCB | 2 |
| 2019 | Pattern recognition analysis on long noncoding RNAs: a tool for prediction in plantsabstractMOTIVATION: Long noncoding RNAs (lncRNAs) correspond to a eukaryotic noncoding RNA class that gained great attention in the past years as a higher layer of regulation for gene expression in cells. There is, however, a lack of specific computational approaches to reliably predict lncRNA in plants, which contrast the variety of prediction tools available for mammalian lncRNAs. This distinction is not that obvious, given that biological features and mechanisms generating lncRNAs in the cell are likely different between animals and plants. Considering this, we present a machine learning analysis and a classifier approach called RNAplonc (https://github.com/TatianneNegri/RNAplonc/) to identify lncRNAs in plants. RESULTS: Our feature selection analysis considered 5468 features, and it used only 16 features to robustly identify lncRNA with the REPTree algorithm. That was the base to create the model and train it with lncRNA and mRNA data from five plant species (thale cress, cucumber, soybean, poplar and Asian rice). After an extensive comparison with other tools largely used in plants (CPC, CPC2, CPAT and PLncPRO), we found that RNAplonc produced more reliable lncRNA predictions from plant transcripts with 87.5% of the best result in eight tests in eight species from the GreeNC database and four independent studies in monocotyledonous (Brachypodium) and eudicotyledonous (Populus and Gossypium) species. Tatianne da Costa Negri, Wonder Alexandre Luz Alves, Pedro Henrique Bugatti, Priscila T. M. Saito, Douglas Silva Domingues, Alexandre Rossi Paschoal |
Briefings Bioinform. | 6 |
| 2019 | mirtronDB: a mirtron knowledge baseabstractMOTIVATION: Mirtrons arise from short introns with atypical cleavage by using the splicing mechanism. In the current literature, there is no repository centralizing and organizing the data available to the public. To fill this gap, we developed mirtronDB, the first knowledge database dedicated to mirtron, and it is available at http://mirtrondb.cp.utfpr.edu.br/. MirtronDB currently contains a total of 1407 mirtron precursors and 2426 mirtron mature sequences in 18 species. RESULTS: Through a user-friendly interface, users can now browse and search mirtrons by organism, organism group, type and name. MirtronDB is a specialized resource that provides free and user-friendly access to knowledge on mirtron data. AVAILABILITY AND IMPLEMENTATION: MirtronDB is available at http://mirtrondb.cp.utfpr.edu.br/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bruno Henrique Ribeiro Da Fonseca, Douglas Silva Domingues, Alexandre Rossi Paschoal |
Bioinform. | 3 |
| 2018 | ceRNAs in plants: computational approaches and associated challenges for target mimic researchabstractThe competing endogenous RNA hypothesis has gained increasing attention as a potential global regulatory mechanism of microRNAs (miRNAs), and as a powerful tool to predict the function of many noncoding RNAs, including miRNAs themselves. Most studies have been focused on animals, although target mimic (TMs) discovery as well as important computational and experimental advances has been developed in plants over the past decade. Thus, our contribution summarizes recent progresses in computational approaches for research of miRNA:TM interactions. We divided this article in three main contributions. First, a general overview of research on TMs in plants is presented with practical descriptions of the available literature, tools, data, databases and computational reports. Second, we describe a common protocol for the computational and experimental analyses of TM. Third, we provide a bioinformatics approach for the prediction of TM motifs potentially cross-targeting both members within the same or from different miRNA families, based on the identification of consensus miRNA-binding sites from known TMs across sequenced genomes, transcriptomes and known miRNAs. This computational approach is promising because, in contrast to animals, miRNA families in plants are large with identical or similar members, several of which are also highly conserved. From the three consensus TM motifs found with our approach: MIM166, MIM171 and MIM159/319, the last one has found strong support on the recent experimental work by Reichel and Millar [Specificity of plant microRNA TMs: cross-targeting of mir159 and mir319. J Plant Physiol 2015;180:45-8]. Finally, we stress the discussion on the major computational and associated experimental challenges that have to be faced in future ceRNA studies. Alexandre Rossi Paschoal, Irma Lozada-Chávez, Douglas Silva Domingues, Peter F. Stadler |
Briefings Bioinform. | 1 |