Diego P. Rubert

dblp:173/8640 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
2since 2021 · last 2022
0000-0002-4131-7309ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 5 first-author · 1 since 2021Theory of computation · 2 · 1 since 2021
YearPublicationVenuePosition
2022 Gene Orthology Inference via Large-Scale Rearrangements for Partially Assembled Genomes
Diego P. Rubert, Marília D. V. Braga
WABI1
2021 Algorithms for Normalized Multiple Sequence Alignments
abstract
Sequence alignment supports numerous tasks in bioinformatics, natural language processing, pattern recognition, social sciences, and other fields. While the alignment of two sequences may be performed swiftly in many applications, the simultaneous alignment of multiple sequences proved to be naturally more intricate. Although most multiple sequence alignment (MSA) formulations are NP-hard, several approaches have been developed, as they can outperform pairwise alignment methods or are necessary for some applications. Taking into account not only similarities but also the lengths of the compared sequences (i.e. normalization) can provide better alignment results than both unnormalized or post-normalized approaches. While some normalized methods have been developed for pairwise sequence alignment, none have been proposed for MSA. This work is a first effort towards the development of normalized methods for MSA. We discuss multiple aspects of normalized multiple sequence alignment (NMSA). We define three new criteria for computing normalized scores when aligning multiple sequences, showing the NP-hardness and exact algorithms for solving the NMSA using those criteria. In addition, we provide approximation algorithms for MSA and NMSA for some classes of scoring matrices.
Eloi Araujo, Luiz C. S. Rozante, Diego P. Rubert, Fábio Viduani Martinez
ISAAC3
2020 Natural Family-Free Genomic Distance
Diego P. Rubert, Fábio Viduani Martinez, Marília D. V. Braga
WABI1
2020 Scalable parallel algorithms for maximum matching and Hamiltonian circuit in convex bipartite graphs
Marco Aurelio Stefanes, Diego P. Rubert, José Soares
Theor. Comput. Sci.2
2018 Computing the family-free DCJ similarity
abstract
BACKGROUND: The genomic similarity is a large-scale measure for comparing two given genomes. In this work we study the (NP-hard) problem of computing the genomic similarity under the DCJ model in a setting that does not assume that the genes of the compared genomes are grouped into gene families. This problem is called family-free DCJ similarity. RESULTS: We propose an exact ILP algorithm to solve the family-free DCJ similarity problem, then we show its APX-hardness and present four combinatorial heuristics with computational experiments comparing their results to the ILP. CONCLUSIONS: We show that the family-free DCJ similarity can be computed in reasonable time, although for larger genomes it is necessary to resort to heuristics. This provides a basis for further studies on the applicability and model refinement of family-free whole genome similarity measures.
Diego P. Rubert, Edna Ayako Hoshino, Marília D. V. Braga, Jens Stoye, Fábio Viduani Martinez
BMC Bioinform.1
2016 A Linear Time Approximation Algorithm for the DCJ Distance for Genomes with Bounded Number of Duplicates
Diego P. Rubert, Pedro Feijão, Marília D. V. Braga, Jens Stoye, Fábio Viduani Martinez
WABI1
2015 SIMBio: Searching and inferring colorful motifs in biological networks
abstract
The study of motifs plays a central role in recognition of relations among components in biological networks such that gene regulation, protein interaction, and metabolic networks. Since these relations are not well-known, motifs inference appears as a way for understanding the principles involved in the relationship between cellular components. On the other hand, motifs search is a basic step for constructing models which represent biological behavior and explain functional and/or structural effects in biological networks. In this work we address the problem of infer all relevant motifs in a biological network. We also provide a solution for searching colorful motifs which can be topological-free or have an acyclic topology. We developed a tool for searching and inferring motifs, named SIMBio, and we implemented sequential and parallel versions. When comparing performance, our experiments have showed that SIMBio is faster than MOTUS for inferring motifs, even in the sequential version. We also compared it to Torque, and SIMBio has found more occurrences of motifs under the same experiments.
Diego P. Rubert, Eloi Araujo, Marco Aurelio Stefanes
BIBE1