Danilo Sipoli Sanches

dblp:202/9057 · DBLP profile ↗
← Back
21ranked-venue papers
4as first author
9since 2021 · last 2024
0000-0002-8972-5221ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-authorSystems, architecture and hardware · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2024 biomapp::chip: large-scale motif analysis
abstract
BACKGROUND: Discovery biological motifs plays a fundamental role in understanding regulatory mechanisms. Computationally, they can be efficiently represented as kmers, making the counting of these elements a critical aspect for ensuring not only the accuracy but also the efficiency of the analytical process. This is particularly useful in scenarios involving large data volumes, such as those generated by the ChIP-seq protocol. Against this backdrop, we introduce BIOMAPP::CHIP, a tool specifically designed to optimize the discovery of biological motifs in large data volumes. RESULTS: We conducted a comprehensive set of comparative tests with state-of-the-art algorithms. Our analyses revealed that BIOMAPP::CHIP outperforms existing approaches in various metrics, excelling both in terms of performance and accuracy. The tests demonstrated a higher detection rate of significant motifs and also greater agility in the execution of the algorithm. Furthermore, the SMT component played a vital role in the system's efficiency, proving to be both agile and accurate in kmer counting, which in turn improved the overall efficacy of our tool. CONCLUSION: BIOMAPP::CHIP represent real advancements in the discovery of biological motifs, particularly in large data volume scenarios, offering a relevant alternative for the analysis of ChIP-seq data and have the potential to boost future research in the field. This software can be found at the following address: (https://github.com/jadermcg/biomapp-chip).
Jader C. Garbelini, Danilo Sipoli Sanches, Aurora T. R. Pozo
BMC Bioinform.2
2024 Simulation of 69 microbial communities indicates sequencing depth and false positives are major drivers of bias in prokaryotic metagenome-assembled genome recovery
abstract
We hypothesize that sample species abundance, sequencing depth, and taxonomic relatedness influence the recovery of metagenome-assembled genomes (MAGs). To test this hypothesis, we assessed MAG recovery in three in silico microbial communities composed of 42 species with the same richness but different sample species abundance, sequencing depth, and taxonomic distribution profiles using three different pipelines for MAG recovery. The pipeline developed by Parks and colleagues (8K) generated the highest number of MAGs and the lowest number of true positives per community profile. The pipeline by Karst and colleagues (DT) showed the most accurate results (~ 92%), outperforming the 8K and Multi-Metagenome pipeline (MM) developed by Albertsen and collaborators. Sequencing depth influenced the accurate recovery of genomes when using the 8K and MM, even with contrasting patterns: the MM pipeline recovered more MAGs found in the original communities when employing sequencing depths up to 60 million reads, while the 8K recovered more true positives in communities sequenced above 60 million reads. DT showed the best species recovery from the same genus, even though close-related species have a low recovery rate in all pipelines. Our results highlight that more bins do not translate to the actual community composition and that sequencing depth plays a role in MAG recovery and increased community resolution. Even low MAG recovery error rates can significantly impact biological inferences. Our data indicates that the scientific community should curate their findings from MAG recovery, especially when asserting novel species or metabolic traits.
Ulisses Nunes da Rocha, Jonas Coelho Kasmanas, Rodolfo Toscan, Danilo Sipoli Sanches, Stefanía Magnúsdóttir, João Pedro Saraiva
PLoS Comput. Biol.4
2024 Three-phase induction motor fault identification using optimization algorithms and intelligent systems
Jacqueline Jordan Guedes, Alessandro Goedtel, Marcelo Favoretto Castoldi, Danilo Sipoli Sanches, Paulo José Amaral Serni, Agnes Fernanda Ferreira Rezende, Gustavo Henrique Bazan, Wesley Angelino de Souza
Soft Comput.4
2024 Binary differential evolution applied to the optimization of the voltage stability margin through the selection of corrective control sets
Rafael Martini Silva, Marcelo Favoretto Castoldi, Alessandro Goedtel, Danilo Sipoli Sanches, Rodrigo A. Ramos
Soft Comput.4
2022 Expectation Maximization based algorithm applied to DNA sequence motif finder
abstract
Finding transcription factor binding sites plays an important role inside bioinformatics. Its correct identification in the promoter regions of co-expressed genes is a crucial step for understanding gene expression mechanisms and creating new drugs and vaccines. The problem of finding motifs consists in seeking conserved patterns in biological datasets of sequences, through using unsupervised learning algorithms. This problem is considered one of the open problems of computational biology, which in its simplest formulation has been proven to be np-hard. Moreover, heuristics and meta-heuristics algorithms have been shown to be very promising in solving combinatorial problems with very large search spaces. In this paper we propose a new algorithm called Biomapp (Biological Motif Application) based on canonical Expectation Maximization that uses the Kullback-Leibler divergence to re-estimate the parameters of statistical model. Furthermore, the algorithm is embedded in an Iterated Local Search, as the local search step and then, we use a hierarchical perturbation operator in order to escape from local optima. The results obtained by this new approach were compared with the state-of-the-art algorithm MEME (Multiple EM Motif Elicitation) showing that Biomapp outperformed this classical technique in several datasets.
Jader C. Garbelini, Danilo Sipoli Sanches, Aurora T. R. Pozo
CEC2
2022 Assessing Batch and Online Learning for Delivery in Full and On Time Predictions
abstract
Improving results by optimizing process execution is one objective of major companies. For these corporations, the main point for achieving better results is the good maintenance of supply chain management. The most important supply chain metric is Delivery in Full and On Time (DIFOT). DIFOT measures how well a supply chain delivers value to the customer. In this work, we bring forward an analysis of DIFOT prediction from large Brazilian food company. More specifically, we compare a batch and online learning algorithm for DIFOT prediction and depict why the latter is suitable for this problem. Furthermore, we report a feature drift analysis to identify whether there are considerable shifts along with the dataset timespan. As a byproduct of this research, we make the dataset used in this analysis publicly available for future research in DIFOT prediction.
Adriano Alves de Lima, Márcie Venâncio Batista, Jean Paul Barddal, Danilo Sipoli Sanches, Luiz Eduardo Soares de Oliveira
IJCNN4
2022 BioAutoML: automated feature engineering and metalearning to predict noncoding RNAs in bacteria
abstract
Recent technological advances have led to an exponential expansion of biological sequence data and extraction of meaningful information through Machine Learning (ML) algorithms. This knowledge has improved the understanding of mechanisms related to several fatal diseases, e.g. Cancer and coronavirus disease 2019, helping to develop innovative solutions, such as CRISPR-based gene editing, coronavirus vaccine and precision medicine. These advances benefit our society and economy, directly impacting people's lives in various areas, such as health care, drug discovery, forensic analysis and food processing. Nevertheless, ML-based approaches to biological data require representative, quantitative and informative features. Many ML algorithms can handle only numerical data, and therefore sequences need to be translated into a numerical feature vector. This process, known as feature extraction, is a fundamental step for developing high-quality ML-based models in bioinformatics, by allowing the feature engineering stage, with design and selection of suitable features. Feature engineering, ML algorithm selection and hyperparameter tuning are often manual and time-consuming processes, requiring extensive domain knowledge. To deal with this problem, we present a new package: BioAutoML. BioAutoML automatically runs an end-to-end ML pipeline, extracting numerical and informative features from biological sequence databases, using the MathFeature package, and automating the feature selection, ML algorithm(s) recommendation and tuning of the selected algorithm(s) hyperparameters, using Automated ML (AutoML). BioAutoML has two components, divided into four modules: (1) automated feature engineering (feature extraction and selection modules) and (2) Metalearning (algorithm recommendation and hyper-parameter tuning modules). We experimentally evaluate BioAutoML in two different scenarios: (i) prediction of the three main classes of noncoding RNAs (ncRNAs) and (ii) prediction of the eight categories of ncRNAs in bacteria, including housekeeping and regulatory types. To assess BioAutoML predictive performance, it is experimentally compared with two other AutoML tools (RECIPE and TPOT). According to the experimental results, BioAutoML can accelerate new studies, reducing the cost of feature engineering processing and either keeping or improving predictive performance. BioAutoML is freely available at https://github.com/Bonidia/BioAutoML.
Robson Bonidia, Anderson P. Avila-Santos, Breno Lívio Silva de Almeida, Peter F. Stadler, Ulisses Nunes da Rocha, Danilo Sipoli Sanches, André C. P. L. F. de Carvalho
Briefings Bioinform.6
2022 MathFeature: feature extraction package for DNA, RNA and protein sequences based on mathematical descriptors
abstract
One of the main challenges in applying machine learning algorithms to biological sequence data is how to numerically represent a sequence in a numeric input vector. Feature extraction techniques capable of extracting numerical information from biological sequences have been reported in the literature. However, many of these techniques are not available in existing packages, such as mathematical descriptors. This paper presents a new package, MathFeature, which implements mathematical descriptors able to extract relevant numerical information from biological sequences, i.e. DNA, RNA and proteins (prediction of structural features along the primary sequence of amino acids). MathFeature makes available 20 numerical feature extraction descriptors based on approaches found in the literature, e.g. multiple numeric mappings, genomic signal processing, chaos game theory, entropy and complex networks. MathFeature also allows the extraction of alternative features, complementing the existing packages. To ensure that our descriptors are robust and to assess their relevance, experimental results are presented in nine case studies. According to these results, the features extracted by MathFeature showed high performance (0.6350-0.9897, accuracy), both applying only mathematical descriptors, but also hybridization with well-known descriptors in the literature. Finally, through MathFeature, we overcame several studies in eight benchmark datasets, exemplifying the robustness and viability of the proposed package. MathFeature has advanced in the area by bringing descriptors not available in other packages, as well as allowing non-experts to use feature extraction techniques.
Robson Bonidia, Douglas Silva Domingues, Danilo Sipoli Sanches, André C. P. L. F. de Carvalho
Briefings Bioinform.3
2021 Feature extraction approaches for biological sequences: a comparative study of mathematical features
abstract
As consequence of the various genomic sequencing projects, an increasing volume of biological sequence data is being produced. Although machine learning algorithms have been successfully applied to a large number of genomic sequence-related problems, the results are largely affected by the type and number of features extracted. This effect has motivated new algorithms and pipeline proposals, mainly involving feature extraction problems, in which extracting significant discriminatory information from a biological set is challenging. Considering this, our work proposes a new study of feature extraction approaches based on mathematical features (numerical mapping with Fourier, entropy and complex networks). As a case study, we analyze long non-coding RNA sequences. Moreover, we separated this work into three studies. First, we assessed our proposal with the most addressed problem in our review, e.g. lncRNA and mRNA; second, we also validate the mathematical features in different classification problems, to predict the class of lncRNA, e.g. circular RNAs sequences; third, we analyze its robustness in scenarios with imbalanced data. The experimental results demonstrated three main contributions: first, an in-depth study of several mathematical features; second, a new feature extraction pipeline; and third, its high performance and robustness for distinct RNA sequence classification. Availability:https://github.com/Bonidia/FeatureExtraction_BiologicalSequences.
Robson Bonidia, Lucas Dias H. Sampaio, Douglas Silva Domingues, Alexandre Rossi Paschoal, Fabrício Martins Lopes, André C. P. L. F. de Carvalho, Danilo Sipoli Sanches
Briefings Bioinform.7
2019 Feature Extraction of Long Non-coding RNAs: A Fourier and Numerical Mapping Approach
Robson Bonidia, Lucas Dias H. Sampaio, Fabrício Martins Lopes, Danilo Sipoli Sanches
CIARP4
2019 Differential evolution applied to line-connected induction motors stator fault identification
Jacqueline Jordan Guedes, Marcelo Favoretto Castoldi, Alessandro Goedtel, Cristiano M. Agulhari, Danilo Sipoli Sanches
Soft Comput.5
2019 A novel multi-objective evolutionary algorithm based on subpopulations for the bi-objective traveling salesman problem
Deyvid Heric Moraes, Danilo Sipoli Sanches, Josimar da Silva Rocha, Jader C. Garbelini, Marcelo Favoretto Castoldi
Soft Comput.2
2018 Sequence motif finder using memetic algorithm
abstract
BACKGROUND: De novo prediction of Transcription Factor Binding Sites (TFBS) using computational methods is a difficult task and it is an important problem in Bioinformatics. The correct recognition of TFBS plays an important role in understanding the mechanisms of gene regulation and helps to develop new drugs. RESULTS: We here present Memetic Framework for Motif Discovery (MFMD), an algorithm that uses semi-greedy constructive heuristics as a local optimizer. In addition, we used a hybridization of the classic genetic algorithm as a global optimizer to refine the solutions initially found. MFMD can find and classify overrepresented patterns in DNA sequences and predict their respective initial positions. MFMD performance was assessed using ChIP-seq data retrieved from the JASPAR site, promoter sequences extracted from the ABS site, and artificially generated synthetic data. The MFMD was evaluated and compared with well-known approaches in the literature, called MEME and Gibbs Motif Sampler, achieving a higher f-score in the most datasets used in this work. CONCLUSIONS: We have developed an approach for detecting motifs in biopolymers sequences. MFMD is a freely available software that can be promising as an alternative to the development of new tools for de novo motif discovery. Its open-source software can be downloaded at https://github.com/jadermcg/mfmd .
Jader C. Garbelini, André Y. Kashiwabara, Danilo Sipoli Sanches
BMC Bioinform.3
2017 Building a better heuristic for the traveling salesman problem: combining edge assembly crossover and partition crossover
abstract
A genetic algorithm using Edge Assemble Crossover (EAX) is one of the best heuristic solvers for large instances of the Traveling Salesman Problem. We propose using Partition Crossover to recombine solutions produced by EAX. Partition Crossover is a powerful deterministic recombination that is highly exploitive. When Partition Crossover decomposes two parents into q recombining components, partition crossover returns the best of 2q reachable offspring. If two parents are locally optimal, all of the offspring are also locally optimal in a hyperplane subspace that contains the two parents. One disadvantage of Partition Crossover, however, is that it cannot generate new edges. By contrast, the EAX operator is highly explorative; it not only inherits edges from parents, it also introduces new edges into the population of a genetic algorithm. Using both EAX and Partition Crossover together produces better performance, with improved exploitation and exploration.
Danilo Sipoli Sanches, L. Darrell Whitley, Renato Tinós
GECCO1
2017 Improving an exact solver for the traveling salesman problem using partition crossover
abstract
The best known exact solver for generating provably optimal solutions to the Traveling Salesman Problem (TSP) is the Concorde algorithm. Concorde uses a branch and bound search strategy, as well as cutting planes to reduce the search space. The first step in using Concorde is to obtain a good initial solution. A good solution can be generated using a heuristic solver outside of Concorde, or Concorde can generate its own initial solution using the Chained Lin Kernighan (LK) algorithm. In this paper, we speed up Concorde by improving the initial solutions produced by Chained LK using Partition Crossover. Partition Crossover is a powerful deterministic recombination operator that is able to tunnel between local optima. In every instance we examined, the addition of recombination resulted in an average speed-up of Concorde, and in the majority of cases, the difference in the runtime costs was statistically significant.
Danilo Sipoli Sanches, L. Darrell Whitley, Renato Tinós
GECCO1
2016 An Evaluation of Different Evolutionary Approaches Applied in the Process of Automatic Transcription of Music Scores into Tablatures
abstract
The problem of converting a music in standard music notation (music sheet) to the alternative notation of guitar tablature is known as transcription. The process of transcription consists of indicating where each note from the original music sheet needs to be played in the guitar, i.e. which string and fret of the guitar that needs to be played to produce a particular note. However, considering that each note can be played in different positions of the guitar fretboard, this is not a straightforward process, and can be classified as a combinatorial optimization problem. For this reason, we have employed a comparative study of different algorithms: A-star, genetic algorithms (GA), genetic algorithms based on subpopulations (GA-SP), ant colony optimization (ACO) and differential evolution (DE). It was also included heuristics based on local search 2-opt and 3-opt in the approaches GA, GA-SP and DE. The experimental results with a dataset of 87 musics indicated that the approaches ACO, GA-SP with 2-opt and GA with 2-opt reached the best performance. Also, the results obtained with each approach were statistically compared using ANOVA test with post hoc Tukey.
Joao Victor Ramos, Andre Stylianos Ramos, Carlos Nascimento Silla Jr., Danilo Sipoli Sanches
ICTAI4
2015 Multi-objective Evolutionary Algorithm with Discrete Differential Mutation Operator for Service Restoration in Large-Scale Distribution Systems
Danilo Sipoli Sanches, Telma Woerle de Lima Soares, João Bosco A. London Jr., Alexandre C. B. Delbem, Ricardo Sérgio Prado, Frederico G. Guimarães
EMO (2)1
2013 Multi-Objective Evolutionary Algorithm with Node-Depth Encoding and Strength Pareto for Service Restoration in Large-Scale Distribution Systems
Marcilyanne Moreira Gois, Danilo Sipoli Sanches, Jean Paulo Martins, João Bosco A. London Jr., Alexandre C. B. Delbem
EMO2
2013 Parallel simultaneous and coordinated tuning of PSSs using Ant Colony Optimization
abstract
Power system controllers are typically designed using trial-and-error techniques, which may require large effort and time from the part of the designer to find a satisfactory solution. This work proposes an algorithm to perform a simultaneous and coordinated tuning of the controllers (PSSs) in an automatic form. To perform this tuning, the algorithm uses an optimization technique based on ant colony metaheuristic. In addition, a parallel structure for the algorithm is proposed to minimize the computation time. Results show satisfactory performance of the tuned controllers, evidencing the effectiveness of the proposed technique. Furthermore, a significant productivity gain can be achieved if the engineer in charge of this design only supervises the automatic process, instead of performing all calculations himself/herself.
Sérgio Carlos Mazucato Júnior, Bruno Leandro Galvao Costa, Marcelo Favoretto Castoldi, Bruno Augusto Angélico, Danilo Sipoli Sanches, Rodrigo A. Ramos
IECON5
2013 Combining subpopulation tables, non-dominated solutions and Strength Pareto of MOEAs to treat service restoration problem in large-scale distribution systems
abstract
The network reconfiguration for service restoration (SR) in distribution systems is a combinatorial complex optimization problem since it involves multiple non-linear constraints and objectives. For large networks, no exact algorithm has found adequate SR plans in real-time. On the other hand, methods combining Multi-objective Evolutionary Algorithms (MOEAs) with the Node-depth encoding (NDE) have shown to be able to efficiently generate adequate SR plans for large distribution systems (with thousands of buses and switches). This paper presents a new method that combining NDE with three MOEAs: (i) NSGA-II; (iii) SPEA 2; and (iii) a MOEA based on subpopulation tables. The idea is to obtain a method that cannot-only obtain adequate SR plans for large scale distribution systems, but can also find plans for small or large networks with similar quality. The proposed method, called MEA2N-STR, explores the space of the objectives solutions better than the other MOEAs with NDE, approximating better the Pareto-optimal front. This statement has been demonstrated by several simulations with DSs ranging from 632 to 1,277 switches.
Danilo Sipoli Sanches, Sérgio Carlos Mazucato Júnior, Marcelo Favoretto Castoldi, Alexandre C. B. Delbem, João Bosco A. London Jr.
IECON1
2009 Modeling Strategy by Adaptive Genetic Algorithm for Production Reactive Scheduling with Simultaneous use of Machines and AGVs
abstract
The problem of production scheduling of manufacturing systems is characterized by the large number of possible solutions. Several researches have been using the Genetic Algorithms (GA) as a search method to solve this problem since these algorithms have the capacity of globally exploring the search space and find good solutions quickly. Since the performance of the GA is directly related to the choice of the parameters of genetic operators, and a bad choice can depreciate the performance, this paper proposes the use of Adaptive Genetic Algorithm to solve this kind of scheduling problem considering the machines and the Automated Guided Vehicles (AGVs). The aim of this paper is to get a good production reactive schedule in order to achieve a good makespan value in a low response obtaining time. The results of this paper were validated in large scenarios and compared with the results of two other approaches. These results are presented and discussed in this paper.
Orides Morandin Jr., Edilson R. R. Kato, Danilo Sipoli Sanches, Bruno Drugowick Muniz
SMC3