EDBT 2026 Demo / reviewers in the wild / expert
Douglas Silva Domingues
dblp:227/0000
· DBLP profile ↗
9ranked-venue papers
0as first author
6since 2021 · last 2026
0000-0002-1290-0853ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | mirtronDB 2.0: enhanced database with novel mirtron discoveriesabstractMOTIVATION: MirtronDB provides a comprehensive and up-to-date resource for advancing mirtron research within RNA biology. Therefore, maintaining a specialized and continuously updated resource for mirtrons is essential to support ongoing discoveries and to serve as a key reference for researchers investigating the roles of mirtrons. RESULTS: Here, we present mirtronDB 2.0, an enhanced version that expands both content and functionality. This version integrates mirtron data published between 2017 and 2025, increasing the number of documented mirtrons across various species. In addition, it incorporates newly predicted mirtrons identified through a robust pipeline that combines advanced bioinformatics and machine learning approaches, with specific coverage of six mammalian species. We have introduced new website features, including an interactive dashboard to enhance usability and facilitate intuitive data exploration. These rigorous updates consolidate mirtronDB as a key resource for mirtron to the RNA biology community. AVAILABILITY AND IMPLEMENTATION: mirtronDB can be found under http://mirtrondb.cp.utfpr.edu.br/. The complete content of Database 2.0 and the source code for the analyses are also freely available in the FigShare repository: https://figshare.com/articles/dataset/MirtronDB_version2/29344775. Fabiana Rodrigues de Góes, Matheus Fujimura Soares, Vitor Gregorio, Bruno Thiago de Lima Nichio, Alisson Gaspar Chiquitto, Flavia Lombardi Lopes, Mark Basham, Douglas Silva Domingues, Alexandre Rossi Paschoal |
Bioinform. | 8 |
| 2024 | Classification of LTR Retrotransposons via Interaction PredictionabstractTransposable Elements (TEs) are genetic sequences that can relocate within the genome, promoting genetic diversity. In eukaryotes, TEs are classified into classes, subclasses, orders, superfamilies, families, and subfamilies. LTR retrotransposons (LTR-RT) constitute an order in this taxonomy. The main objective of this study is to investigate the classification of LTR retrotransposons at the superfamily level. Predictive Bi-Clustering Trees (PBCTs) were used to predict interactions between LTR-RT sequences and conserved protein domains to achieve this. Two datasets were used to investigate the relationships among different superfamilies. The first dataset contained LTR retrotransposon sequences assigned to Copia, Gypsy, and Bel-Pao superfamilies, while the second dataset included consensus sequences of the conserved domains for each superfamily. Thus, the PBCT decision tree tests could relate to the LTR-RT sequence and conserved domain attributes. In the classification process, interaction is interpreted as either the presence or absence of a domain in a given LTR-RT sequence. The sequence is then classified into the superfamily with the most predicted domains. Precision-recall curves were adopted as evaluation metrics for the method, and its performance was compared to some of the most commonly used models in the task of transposable element classification. The experiments conducted on D. melanogaster and A. thaliana showed that PBCTs are promising and comparable to other methods, especially in classifying the Gypsy superfamily. Silvana C. S. Cardoso, Douglas Silva Domingues, Alexandre Rossi Paschoal, Carlos Fischer, Ricardo Cerri |
CIBCB | 2 |
| 2022 | Impact of sequencing technologies on long non-coding RNA computational identificationabstractThe correct annotation of non-coding RNAs, especially long non-coding RNAs (lncRNAs), is still a critial challenge in genome analyses due to their highly heterogeneous characteristics. Due to this heterogeneity, transcriptome data sources can be an important factor that might affect lncRNA annotation quality. Long-read technologies now bring the potential to improve the quality of transcriptome annotation, specially in genome entities that are not “classic” coding genes. However, there is a gap regarding benchmarking studies that test if the direct use of lncRNA predictors in long-reads makes more precise identification of these transcripts. Considering that lncRNA identification tools were not trained with these reads, our study want to address: how is the performance of these tools? Are they also able to efficiently identify lncRNAs? For this, we used short and long-read data from human and selected plants transcriptomes to test our questions. We can provide evidence of where and how to make potential better approaches for the lncRNA annotation by understanding these issues. Alisson Gaspar Chiquitto, Lucas Otávio L. Silva, Liliane S. Oliveira, Douglas Silva Domingues, Alexandre Rossi Paschoal |
BIBM | 4 |
| 2022 | MathFeature: feature extraction package for DNA, RNA and protein sequences based on mathematical descriptorsabstractOne of the main challenges in applying machine learning algorithms to biological sequence data is how to numerically represent a sequence in a numeric input vector. Feature extraction techniques capable of extracting numerical information from biological sequences have been reported in the literature. However, many of these techniques are not available in existing packages, such as mathematical descriptors. This paper presents a new package, MathFeature, which implements mathematical descriptors able to extract relevant numerical information from biological sequences, i.e. DNA, RNA and proteins (prediction of structural features along the primary sequence of amino acids). MathFeature makes available 20 numerical feature extraction descriptors based on approaches found in the literature, e.g. multiple numeric mappings, genomic signal processing, chaos game theory, entropy and complex networks. MathFeature also allows the extraction of alternative features, complementing the existing packages. To ensure that our descriptors are robust and to assess their relevance, experimental results are presented in nine case studies. According to these results, the features extracted by MathFeature showed high performance (0.6350-0.9897, accuracy), both applying only mathematical descriptors, but also hybridization with well-known descriptors in the literature. Finally, through MathFeature, we overcame several studies in eight benchmark datasets, exemplifying the robustness and viability of the proposed package. MathFeature has advanced in the area by bringing descriptors not available in other packages, as well as allowing non-experts to use feature extraction techniques. Robson Bonidia, Douglas Silva Domingues, Danilo Sipoli Sanches, André C. P. L. F. de Carvalho |
Briefings Bioinform. | 2 |
| 2021 | Feature extraction approaches for biological sequences: a comparative study of mathematical featuresabstractAs consequence of the various genomic sequencing projects, an increasing volume of biological sequence data is being produced. Although machine learning algorithms have been successfully applied to a large number of genomic sequence-related problems, the results are largely affected by the type and number of features extracted. This effect has motivated new algorithms and pipeline proposals, mainly involving feature extraction problems, in which extracting significant discriminatory information from a biological set is challenging. Considering this, our work proposes a new study of feature extraction approaches based on mathematical features (numerical mapping with Fourier, entropy and complex networks). As a case study, we analyze long non-coding RNA sequences. Moreover, we separated this work into three studies. First, we assessed our proposal with the most addressed problem in our review, e.g. lncRNA and mRNA; second, we also validate the mathematical features in different classification problems, to predict the class of lncRNA, e.g. circular RNAs sequences; third, we analyze its robustness in scenarios with imbalanced data. The experimental results demonstrated three main contributions: first, an in-depth study of several mathematical features; second, a new feature extraction pipeline; and third, its high performance and robustness for distinct RNA sequence classification. Availability:https://github.com/Bonidia/FeatureExtraction_BiologicalSequences. Robson Bonidia, Lucas Dias H. Sampaio, Douglas Silva Domingues, Alexandre Rossi Paschoal, Fabrício Martins Lopes, André C. P. L. F. de Carvalho, Danilo Sipoli Sanches |
Briefings Bioinform. | 3 |
| 2021 | TERL: classification of transposable elements by convolutional neural networksabstractTransposable elements (TEs) are the most represented sequences occurring in eukaryotic genomes. Few methods provide the classification of these sequences into deeper levels, such as superfamily level, which could provide useful and detailed information about these sequences. Most methods that classify TE sequences use handcrafted features such as k-mers and homology-based search, which could be inefficient for classifying non-homologous sequences. Here we propose an approach, called transposable elements pepresentation learner (TERL), that preprocesses and transforms one-dimensional sequences into two-dimensional space data (i.e., image-like data of the sequences) and apply it to deep convolutional neural networks. This classification method tries to learn the best representation of the input data to classify it correctly. We have conducted six experiments to test the performance of TERL against other methods. Our approach obtained macro mean accuracies and F1-score of 96.4% and 85.8% for superfamilies and 95.7% and 91.5% for the order sequences from RepBase, respectively. We have also obtained macro mean accuracies and F1-score of 95.0% and 70.6% for sequences from seven databases into superfamily level and 89.3% and 73.9% for the order level, respectively. We surpassed accuracy, recall and specificity obtained by other methods on the experiment with the classification of order level sequences from seven databases and surpassed by far the time elapsed of any other method for all experiments. Therefore, TERL can learn how to predict any hierarchical level of the TEs classification system and is about 20 times and three orders of magnitude faster than TEclass and PASTEC, respectively https://github.com/muriloHoracio/TERL. Contact:[email protected]. Murilo Horacio Pereira da Cruz, Douglas Silva Domingues, Priscila T. M. Saito, Alexandre Rossi Paschoal, Pedro Henrique Bugatti |
Briefings Bioinform. | 2 |
| 2019 | Pattern recognition analysis on long noncoding RNAs: a tool for prediction in plantsabstractMOTIVATION: Long noncoding RNAs (lncRNAs) correspond to a eukaryotic noncoding RNA class that gained great attention in the past years as a higher layer of regulation for gene expression in cells. There is, however, a lack of specific computational approaches to reliably predict lncRNA in plants, which contrast the variety of prediction tools available for mammalian lncRNAs. This distinction is not that obvious, given that biological features and mechanisms generating lncRNAs in the cell are likely different between animals and plants. Considering this, we present a machine learning analysis and a classifier approach called RNAplonc (https://github.com/TatianneNegri/RNAplonc/) to identify lncRNAs in plants. RESULTS: Our feature selection analysis considered 5468 features, and it used only 16 features to robustly identify lncRNA with the REPTree algorithm. That was the base to create the model and train it with lncRNA and mRNA data from five plant species (thale cress, cucumber, soybean, poplar and Asian rice). After an extensive comparison with other tools largely used in plants (CPC, CPC2, CPAT and PLncPRO), we found that RNAplonc produced more reliable lncRNA predictions from plant transcripts with 87.5% of the best result in eight tests in eight species from the GreeNC database and four independent studies in monocotyledonous (Brachypodium) and eudicotyledonous (Populus and Gossypium) species. Tatianne da Costa Negri, Wonder Alexandre Luz Alves, Pedro Henrique Bugatti, Priscila T. M. Saito, Douglas Silva Domingues, Alexandre Rossi Paschoal |
Briefings Bioinform. | 5 |
| 2019 | mirtronDB: a mirtron knowledge baseabstractMOTIVATION: Mirtrons arise from short introns with atypical cleavage by using the splicing mechanism. In the current literature, there is no repository centralizing and organizing the data available to the public. To fill this gap, we developed mirtronDB, the first knowledge database dedicated to mirtron, and it is available at http://mirtrondb.cp.utfpr.edu.br/. MirtronDB currently contains a total of 1407 mirtron precursors and 2426 mirtron mature sequences in 18 species. RESULTS: Through a user-friendly interface, users can now browse and search mirtrons by organism, organism group, type and name. MirtronDB is a specialized resource that provides free and user-friendly access to knowledge on mirtron data. AVAILABILITY AND IMPLEMENTATION: MirtronDB is available at http://mirtrondb.cp.utfpr.edu.br/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bruno Henrique Ribeiro Da Fonseca, Douglas Silva Domingues, Alexandre Rossi Paschoal |
Bioinform. | 2 |
| 2018 | ceRNAs in plants: computational approaches and associated challenges for target mimic researchabstractThe competing endogenous RNA hypothesis has gained increasing attention as a potential global regulatory mechanism of microRNAs (miRNAs), and as a powerful tool to predict the function of many noncoding RNAs, including miRNAs themselves. Most studies have been focused on animals, although target mimic (TMs) discovery as well as important computational and experimental advances has been developed in plants over the past decade. Thus, our contribution summarizes recent progresses in computational approaches for research of miRNA:TM interactions. We divided this article in three main contributions. First, a general overview of research on TMs in plants is presented with practical descriptions of the available literature, tools, data, databases and computational reports. Second, we describe a common protocol for the computational and experimental analyses of TM. Third, we provide a bioinformatics approach for the prediction of TM motifs potentially cross-targeting both members within the same or from different miRNA families, based on the identification of consensus miRNA-binding sites from known TMs across sequenced genomes, transcriptomes and known miRNAs. This computational approach is promising because, in contrast to animals, miRNA families in plants are large with identical or similar members, several of which are also highly conserved. From the three consensus TM motifs found with our approach: MIM166, MIM171 and MIM159/319, the last one has found strong support on the recent experimental work by Reichel and Millar [Specificity of plant microRNA TMs: cross-targeting of mir159 and mir319. J Plant Physiol 2015;180:45-8]. Finally, we stress the discussion on the major computational and associated experimental challenges that have to be faced in future ceRNA studies. Alexandre Rossi Paschoal, Irma Lozada-Chávez, Douglas Silva Domingues, Peter F. Stadler |
Briefings Bioinform. | 3 |