EDBT 2026 Demo / reviewers in the wild / expert
Diego Furtado Silva
dblp:124/2376
· DBLP profile ↗
35ranked-venue papers
7as first author
11since 2021 · last 2025
0000-0002-5184-9413ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 4 first-author · 9 since 2021Databases, data management, data science and information retrieval · 17 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Local LLM Ensembles for Zero-Shot Portuguese Named Entity Recognition
João Lucas Luz Lima Sarcinelli, Diego Furtado Silva |
CIARP | 2 |
| 2025 | Match: A Maximum-Likelihood Approach for Classification under Label ShiftabstractMachine learning models often suffer from performance degradation when dealing with class distributions that differ from the training distribution, a scenario commonly referred to as label shift. Addressing this challenge, this paper introduces Match, a novel adjustment approach that maximizes the likelihood of predicted probabilities under class prevalence constraints. Unlike existing methods such as retraining with instance re-weighting and the Bayes update rule, Match ensures that the adjusted class distribution aligns precisely with the prevalence estimates from quantifiers. By formulating the adjustment process as a binary integer linear optimization problem, Match benefits from efficient mixed-integer solvers. Extensive experiments demonstrate that Match outperforms the state-of-the-art in classifier adjustment with statistical significance, particularly in handling scenarios with imbalanced distributions. Zahra Donyavi, Feiyu Li, Yunrui Zhang, Diego Furtado Silva, Gustavo Batista |
KDD (2) | 4 |
| 2025 | One-class graph autoencoder: A new end-to-end, low-dimensional, and interpretable approach for node classification
Marcos P. S. Gôlo, José Gilberto Barbosa de Medeiros Júnior, Diego Furtado Silva, Ricardo M. Marcacini |
Inf. Sci. | 3 |
| 2024 | Unsupervised feature based algorithms for time series extrinsic regressionabstractAbstract Time Series Extrinsic Regression (TSER) involves using a set of training time series to form a predictive model of a continuous response variable that is not directly related to the regressor series. The TSER archive for comparing algorithms was released in 2022 with 19 problems. We increase the size of this archive to 63 problems and reproduce the previous comparison of baseline algorithms. We then extend the comparison to include a wider range of standard regressors and the latest versions of TSER models used in the previous study. We show that none of the previously evaluated regressors can outperform a regression adaptation of a standard classifier, rotation forest. We introduce two new TSER algorithms developed from related work in time series classification. FreshPRINCE is a pipeline estimator consisting of a transform into a wide range of summary features followed by a rotation forest regressor. DrCIF is a tree ensemble that creates features from summary statistics over random intervals. Our study demonstrates that both algorithms, along with InceptionTime, exhibit significantly better performance compared to the other 18 regressors tested. More importantly, DrCIF is the only one that significantly outperforms a standard rotation forest regressor. David Guijo-Rubio, Matthew Middlehurst, Guilherme Arcencio, Diego Furtado Silva, Anthony J. Bagnall |
Data Min. Knowl. Discov. | 4 |
| 2024 | Artist Similarity Based on Heterogeneous Graph Neural NetworksabstractMusic streaming platforms rely on recommending similar artists to maintain user engagement, with artists benefiting from these suggestions to boost their popularity. Another important feature is music information retrieval, allowing users to explore new content. In both scenarios, performance depends on how to compute the similarity between musical content. This is a challenging process since musical data is inherently multimodal, containing textual and audio data. We propose a novel graph-based artist representation that integrates audio, lyrics features, and artist relations. Thus, a multimodal representation on a heterogeneous graph is proposed, along with a network regularization process followed by a GNN model to aggregate multimodal information into a more robust unified representation. The proposed method explores this final multimodal representation for the task of artist similarity as a link prediction problem. Our method introduces a new importance matrix to emphasize related artists in this multimodal space. We compare our approach with other strong baselines based on combining input features, importance matrix construction, and GNN models. Experimental results highlight the superiority of multimodal representation through the transfer learning process and the value of the importance matrix in enhancing GNN models for artist similarity. Angelo Cesar Mendes da Silva, Diego Furtado Silva, Ricardo M. Marcacini |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | A New Time Series Framework for Forest Fire Risk Forecasting and ClassificationabstractThere's an increasing concern about the occurrence and spread of forest fires across the globe, as they contribute to greenhouse gas emissions and play a major influential role in economics and public health. Thus, there's a need for accurate methods to predict and classify forest fire risk. The main known forest fire risk indexes have limitations, such as not taking into account the unique characteristics of the biome in study, and not being able to predict forest fire risk for a given number of days in the future. This last aspect, in particular, is of utmost relevance. Addressing it allows for coordinated planning and action by proper authorities with adequate anticipation. Aiming to solve this problem, we present a new framework that applies Machine Learning methods for: (1) climatic variables forecasting; and (2) forest fire risk classification. For the first objective, different time series forecasting algorithms were tested. The forecasted variables are then used as input for the second objective, for which different classification algorithms were also tested. We evaluated our proposal using Brazilian Pantanal regional biome data from 1999 to 2019, where climatic variables were collected from ground meteorological stations, and fire occurrences (hotspots) were obtained from satellite images. The experiments considered 4 climatic variables and 5 forest fire risk classes. The results were evaluated based on the average correlation between (i) the prediction of forest fire risk classes and (ii) the observation of hotspots. Our proposal proved to be better or competitive with the main forest fire risk indexes, with the advantage of predicting fire risk for a given number of days in the future. Bruna Zamith Santos, Balbina Maria Araujo Soriano, Marcelo Gonçalves Narciso, Diego Furtado Silva, Ricardo Cerri |
IJCNN | 4 |
| 2022 | Multimodal representation learning over heterogeneous networks for tag-based music retrieval
Angelo Cesar Mendes da Silva, Diego Furtado Silva, Ricardo M. Marcacini |
Expert Syst. Appl. | 2 |
| 2021 | Financial Time Series Forecasting Enriched with Textual InformationabstractThe ability to extract knowledge and forecast stock trends is crucial to mitigate investors’ risks and uncertainties in the market. The stock trend is affected by non-linearity, complexity, noise, and especially the surrounding news. External factors such as daily news became one of the investors’ primary resources for buying or selling assets. However, this kind of information appears very fast. There are thousands of news generated by different web sources, taking a long time to analyze them, causing significant losses for investors due to late decisions. Although recent contextual language models have transformed the area of natural language processing, models to make predictions using news that influence stock values still face barriers such as unlabeled data and class imbalance. This paper proposes a hybrid methodology that enriches the time series forecasting considering textual knowledge extracted from sites without a widely annotated corpus. We show that the proposed method can improve forecasting using an empirical evaluation of Bitcoin prices prediction. Lord Flaubert Steve Ataucuri Cruz, Diego Furtado Silva |
ICMLA | 2 |
| 2021 | CurL-AutoML: Curriculum Learning-based AutoMLabstractAutoML aims to find the best Machine Learning (ML) pipeline in a complex and high-dimensional search space by evaluating multiple algorithm configurations. However, training multiple ML algorithms is time-consuming, and as AutoML tools are frequently time-constrained, the exploration of the search space may find sub-optimal results. In this work, we explore the application of curriculum learning techniques to overcome this limitation. Curriculum and anti-curriculum learning have improved model performance and accelerated the training process on previous empirical investigations using optimization-based models by ordering examples during model training based on their difficulty. We apply and compare curriculum strategies on an AutoML system to accelerate the search space exploration and find good-performing machine learning pipelines efficiently. The results indicate that AutoML can benefit from a curriculum strategy. Furthermore, in most of the evaluated scenarios, the curriculum strategies led to better classification results. Lucas Nildaimon dos Santos Silva, Lucas Cardoso Silva, Fernando Rezende Zagatti, Bruno Silva Sette, Helena de Medeiros Caseli, Daniel Lucrédio, Diego Furtado Silva |
ICMLA | 7 |
| 2021 | MetaPrep: Data preparation pipelines recommendation via meta-learningabstractData preparation is a mandatory phase in the machine learning pipeline. The goal of data preparation is to convert noisy and disordered data into refined data that can be used by the algorithms. However, data preparation is time-consuming and requires specialized knowledge about the data and algorithms. Therefore, automating data preparation is essential to decrease the effort made by data scientists to develop satisfactory models. Despite its relevance, current AutoML platforms disregard or make simple hardcoded data preparation pipelines. Trying to fill this gap, we present a meta-learning-based recommendation system for data preparation. Our system recommends five pipelines, ranked by their relevance, making it useful for users with varying degrees of experience. Using the top-1 pipeline we demonstrated that our proposal allows a better performance of an AutoML system. Furthermore, the accuracy rates of our method were comparable to those achieved by a reinforcement-learning-based algorithm with the same goal, but it was up to two orders of magnitude faster. Moreover, we tested our method in a real-world application and evaluated its benefits and limitations in this scenario. Fernando Rezende Zagatti, Lucas Cardoso Silva, Lucas Nildaimon dos Santos Silva, Bruno Silva Sette, Helena de Medeiros Caseli, Daniel Lucrédio, Diego Furtado Silva |
ICMLA | 7 |
| 2021 | Committee of NAS-based modelsabstractNetwork Architecture Search (NAS) has achieved impressive results and generated models comparable with humans' classifications. Automating the definition of a neural architecture reduces the need for expert work efforts and mitigates human bias from architecture design. NAS techniques usually consist of an algorithm to search for the best architecture in a predetermined space of parameters or functions. Due to the number of deep neural architectures' parameters, this search space includes millions of parameters, which makes NAS a cost procedure and may lead the search to overfit the training set. To reduce NAS search spaces' complexity and still obtain competitive results, we propose CoNAS, a committee of NAS-based models, by restricting the search spaces to perform Differentiable ARchiTecture Search (DARTS). Our results point to improved accuracy over DARTS on CIFAR-10, training the networks from scratch. and Imagnette, using a transfer learning approach. Bruno Silva Sette, Lucas Cardoso Silva, Fernando Rezende Zagatti, Lucas Nildaimon dos Santos Silva, Daniel Lucrédio, Helena de Medeiros Caseli, Diego Furtado Silva |
IJCNN | 7 |
| 2020 | On Convolutional Autoencoders to Speed Up Similarity-Based Time Series MiningabstractTime series represent a type of data that is increasingly present in research and industry applications. The most commonly used approach to obtain knowledge from these data is similarity-based data mining algorithms. However, in large volumes of data, applying these algorithms may be infeasible. Therefore, several techniques are proposed in the literature to accelerate the distance calculation between time series, such as algorithms' adaptations, indexing, and approximations. Recent results show that the union of techniques in different categories usually improves efficiency in time series mining tasks by similarity. In this way, we evaluate the dimensionality reduction through convolutional autoencoders to speed up time series distance calculation. To this end, we propose an offline training phase, with data previously observed in other domains, before applying the autoencoders. Our proposal is orthogonal to state-of-the-art tools so that we can use it as a complementary technique. We show that our proposal can lead similarity-based algorithms to execute up to two orders of magnitude faster than these tools alone, without loss of quality in the results obtained through all-pairwise distance calculation, similarity search, and motif discovery. Yuri Gabriel Aragão da Silva, Diego Furtado Silva |
IEEE BigData | 2 |
| 2020 | Towards logical association rule mining on ontology-based semantic trajectoriesabstractMobility patterns have been investigated from multiple point-of-views, and different trajectory data mining paradigms have been proposed in the literature. Recent Semantic Trajectory representations rely on ontologies or RDF triples to represent different semantics, interlink data, and retrieve information. Nonetheless, current Trajectory Data Mining approaches use traditional transaction-based techniques that are not compatible with the relational nature of ontology-based representations. In this paper, we tackle the task of association rules mining by borrowing from the Knowledge Base Refinement field, the state-of-the-art AMIE 3 algorithm. After mining logical rules from a Foursquare dataset, we discuss the issues of applying off-the-shelf mining algorithms and discuss opportunities to develop a domain-tailored approach. Antonio Carlos Falcão Petri, Diego Furtado Silva |
ICMLA | 2 |
| 2020 | Benchmarking Machine Learning Solutions in ProductionabstractMachine learning (ML) is becoming critical to many businesses. Keeping an ML solution online and responding is therefore a necessity, and is part of the MLOps (Machine Learning operationalization) movement. One aspect for this process is monitoring not only prediction quality, but also system resources. This is important to correctly provide the necessary infrastructure, either using a fully-managed cloud platform or a local solution. This is not a difficult task, as there are many tools available. However, it requires some planning and knowledge about what to monitor. Also, many ML professionals are not experts in system operations and may not have the skills to easily setup a monitoring and benchmarking environment. In the spirit of MLOps, this paper presents an approach, based on a simple API and set of tools, to monitor ML solutions. The approach was tested with 9 different solutions. The results indicate that the approach can deliver useful information to help in decision making, proper resource provision and operation of ML systems. Lucas Cardoso Silva, Fernando Rezende Zagatti, Bruno Silva Sette, Lucas Nildaimon dos Santos Silva, Daniel Lucrédio, Diego Furtado Silva, Helena de Medeiros Caseli |
ICMLA | 6 |
| 2020 | The Swiss army knife of time series data mining: ten useful things you can do with the matrix profile and ten lines of code
Yan Zhu 0014, Shaghayegh Gharghabi, Diego Furtado Silva, Hoang Anh Dau, Chin-Chia Michael Yeh, Nader Shakibay Senobari, Abdulaziz Almaslukh, Kaveh Kamgar, Zachary Schall-Zimmerman, Gareth J. Funning, Abdullah Mueen, Eamonn J. Keogh |
Data Min. Knowl. Discov. | 3 |
| 2019 | Fast Similarity Matrix Profile for Music Analysis and ExplorationabstractMost algorithms for music data mining and retrieval analyze the similarity between feature sets extracted from the raw audio. A conventional approach to assess similarities within or between recordings is to create similarity matrices. However, this method requires quadratic space for each comparison and typically requires costly post-processing of the matrix. We have recently proposed SiMPle, a powerful representation based on subsequence similarity join, which is applicable in several music analysis tasks. In this paper, we propose SiMPle-Fast a highly efficient method for exact computation of SiMPle that is up to one order of magnitude faster than SiMPle. Furthermore, we demonstrate the utility of SiMPle-Fast in cover music recognition and thumbnailing tasks and show that our method is significantly faster and more accurate than the state-of-the-art. Diego Furtado Silva, Chin-Chia Michael Yeh, Yan Zhu 0014, Gustavo Batista, Eamonn J. Keogh |
IEEE Trans. Multim. | 1 |
| 2018 | Elastic Time Series Motifs and DiscordsabstractThe recent proposal of the Matrix Profile (MP) has brought the attention of the time series community to the usefulness and versatility of the similarity joins. This primitive has numerous applications including the discovery of time series motifs and discords. However, the original MP algorithm has two prominent limitations: the algorithm only works for Euclidean distance (ED) and it is sensitive to the subsequences length. Is this work, we extend the MP algorithm to overcome both limitations. We use a recently proposed variant of Dynamic Time Warping (DTW), the Prefix and Suffix Invariant DTW (PSI-DTW) distance. The PSI-DTW allows invariance to warp and spurious endpoints caused by segmenting subsequences and has a side-effect of supporting the match of subsequences with different lengths. Besides, we propose a suite of simple methods to speed up the MP calculation, making it more than one order of magnitude faster than a straightforward implementation and providing an anytime feature. We show that using PSI-DTW avoids false positives and false dismissals commonly observed by applying ED, improving the time series motifs and discords discovery in several application domains. Diego Furtado Silva, Gustavo Batista |
ICMLA | 1 |
| 2018 | Classifying and Counting with Recurrent ContextsabstractMany real-world applications in the batch and data stream settings with data shift pose restrictions to the access to class labels after the deployment of a classification or quantification model. However, a significant portion of the data stream literature assumes that actual labels are instantaneously available after issuing their corresponding classifications. In this paper, we explore a different set of assumptions without relying on the availability of class labels. We assume that, although the distribution of the data may change over time, it will switch between one of a handful of well-known distributions. Still, we allow the proportions of the classes to vary. In these conditions, we propose the first method that can accurately identify the correct context of data samples and simultaneously estimate the proportion of the positive class. This estimate can be further used to adjust a classification decision threshold and improve classification accuracy. Finally, the method is very efficient regarding time and memory requirements, fitting data stream applications. Denis Moreira dos Reis, André Gustavo Maletzke, Diego Furtado Silva, Gustavo Batista |
KDD | 3 |
| 2018 | Optimizing dynamic time warping's window width for time series data mining applications
Hoang Anh Dau, Diego Furtado Silva, François Petitjean, Germain Forestier, Anthony J. Bagnall, Abdullah Mueen, Eamonn J. Keogh |
Data Min. Knowl. Discov. | 2 |
| 2018 | Speeding up similarity search under dynamic time warping by pruning unpromising alignments
Diego Furtado Silva, Rafael Giusti, Eamonn J. Keogh, Gustavo Batista |
Data Min. Knowl. Discov. | 1 |
| 2018 | Time series joins, motifs, discords and shapelets: a unifying view that exploits the matrix profile
Chin-Chia Michael Yeh, Yan Zhu 0014, Liudmila Ulanova, Nurjahan Begum, Yifei Ding, Hoang Anh Dau, Zachary Schall-Zimmerman, Diego Furtado Silva, Abdullah Mueen, Eamonn J. Keogh |
Data Min. Knowl. Discov. | 8 |
| 2017 | Judicious setting of Dynamic Time Warping's window width allows more accurate classification of time seriesabstractWhile the Dynamic Time Warping (DTW) — based Nearest-Neighbor Classification algorithm is regarded as a strong baseline for time series classification, in recent years there has been a plethora of algorithms that have claimed to be able to improve upon its accuracy in the general case. Many of these proposed ideas sacrifice the simplicity of implementation that DTW-based classifiers offer for rather modest gains. Nevertheless, there are clearly times when even a small improvement could make a large difference in an important medical or financial domain. In this work, we make an unexpected claim; an underappreciated “low hanging fruit” in optimizing DTW's performance can produce improvements that make it an even stronger baseline, closing most or all the improvement gap of the more sophisticated methods. We show that the method currently used to learn DTW's only parameter, the maximum amount of warping allowed, is likely to give the wrong answer for small training sets. We introduce a simple method to mitigate the small training set issue by creating synthetic exemplars to help learn the parameter. We evaluate our ideas on the UCR Time Series Archive and a case study in fall classification, and demonstrate that our algorithm produces significant improvement in classification accuracy. Hoang Anh Dau, Diego Furtado Silva, François Petitjean, Germain Forestier, Anthony J. Bagnall, Eamonn J. Keogh |
IEEE BigData | 2 |
| 2016 | Prefix and Suffix Invariant Dynamic Time WarpingabstractWhile there exist a plethora of classification algorithms for most data types, there is an increasing acceptance that the unique properties of time series mean that the combination of nearest neighbor classifiers and Dynamic Time Warping (DTW) is very competitive across a host of domains, from medicine to astronomy to environmental sensors. While there has been significant progress in improving the efficiency and effectiveness of DTW in recent years, in this work we demonstrate that an underappreciated issue can significantly degrade the accuracy of DTW in real-world deployments. This issue has probably escaped the attention of the very active time series research community because of its reliance on static highly contrived benchmark datasets, rather than real world dynamic datasets where the problem tends to manifest itself. In essence, the issue is that DTW's eponymous invariance to warping is only true for the main "body" of the two time series being compared. However, for the "head" and "tail" of the time series, the DTW algorithm affords no warping invariance. The effect of this is that tiny differences at the beginning or end of the time series (which may be either consequential or simply the result of poor "cropping") will tend to contribute disproportionally to the estimated similarity, producing incorrect classifications. In this work, we show that this effect is real, and reduces the performance of the algorithm. We further show that we can fix the issue with a subtle redesign of the DTW algorithm, and that we can learn an appropriate setting for the extra parameter we introduced. We further demonstrate that our generalization is amiable to all the optimizations that make DTW tractable for large datasets. Diego Furtado Silva, Gustavo Batista, Eamonn J. Keogh |
ICDM | 1 |
| 2016 | Matrix Profile I: All Pairs Similarity Joins for Time Series: A Unifying View That Includes Motifs, Discords and ShapeletsabstractThe all-pairs-similarity-search (or similarity join) problem has been extensively studied for text and a handful of other datatypes. However, surprisingly little progress has been made on similarity joins for time series subsequences. The lack of progress probably stems from the daunting nature of the problem. For even modest sized datasets the obvious nested-loop algorithm can take months, and the typical speed-up techniques in this domain (i.e., indexing, lower-bounding, triangular-inequality pruning and early abandoning) at best produce one or two orders of magnitude speedup. In this work we introduce a novel scalable algorithm for time series subsequence all-pairs-similarity-search. For exceptionally large datasets, the algorithm can be trivially cast as an anytime algorithm and produce high-quality approximate solutions in reasonable time. The exact similarity join algorithm computes the answer to the time series motif and time series discord problem as a side-effect, and our algorithm incidentally provides the fastest known algorithm for both these extensively-studied problems. We demonstrate the utility of our ideas for two time series data mining problems, including motif discovery and novelty discovery. Chin-Chia Michael Yeh, Yan Zhu 0014, Liudmila Ulanova, Nurjahan Begum, Yifei Ding, Hoang Anh Dau, Diego Furtado Silva, Abdullah Mueen, Eamonn J. Keogh |
ICDM | 7 |
| 2016 | Improved Time Series Classification with Representation Diversity and SVMabstractTime series classification is an important task in data mining that has been traditionally addressed with the use of similarity-based classifiers. The 1-NN DTW is typically considered the most accurate model for temporal data. Nevertheless, some authors have recently proposed ingenious alternatives to the 1-NN DTW by using diversity of time series representation or by using DTW for feature extraction. In this paper, we explore diversity of time series representations and distance functions to obtain distance features, which in turn are used to train an SVM model. We argue that the scientific community has largely neglected a vast body of unconventional distance functions, and we present empirical evidence that distance features are better than the 1-NN DTW with respect to classification accuracy. Rafael Giusti, Diego Furtado Silva, Gustavo Batista |
ICMLA | 2 |
| 2016 | Speeding Up All-Pairwise Dynamic Time Warping Matrix CalculationabstractDynamic Time Warping (DTW) is certainly the most relevant distance for time series analysis. However, its quadratic time complexity may hamper its use, mainly in the analysis of large time series data. All the recent advances in speeding up the exact DTW calculation are confined to similarity search. However, there is a significant number of important algorithms including clustering and classification that require the pairwise distance matrix for all time series objects. The only techniques available to deal with this issue are constraint bands and DTW approximations. In this paper, we propose the first exact approach for speeding up the all-pairwise DTW matrix calculation. Our method is exact and may be applied in conjunction with constraint bands. We demonstrate that our algorithm reduces the runtime in approximately 50% on average and up to one order of magnitude in some datasets. Diego Furtado Silva, Gustavo Batista |
SDM | 1 |
| 2015 | Classification of Evolving Data Streams with Infinitely Delayed LabelsabstractThe majority of evolving data streams classification algorithms assume that the actual labels of the predicted examples are readily available without any time delay just after a prediction is made. However, given the high label costs, dependence of an expert, limitations in data transmission or even restrictions imposed by the problem's nature, there is a large number of real-world applications in which the availability of actual labels is infinitely delayed (never available). In these cases, it is necessary the use of algorithms that does not follow the traditional process of monitoring the error rate to detect changes in data distribution and uses the most recent labeled data to update the classification model. In this paper, we propose the method MClassification to classify evolving data streams with infinitely delayed labels. Our method is inspired on the use of Micro-Cluster representation from online clustering algorithms. Considering the presence of incremental drifts, our approach uses a distance-based strategy to maintain the Micro-Clusters' positions updated. An evaluation in several synthetic and real data shows that MClassification achieves competitive accuracy results to state-of-the-art methods and adequate computational cost. The main advantage of the proposed method is the absence of critical parameters that require user's prior knowledge, as occurs with rival methods. Vinícius M. A. de Souza, Diego Furtado Silva, Gustavo Batista, João Gama 0001 |
ICMLA | 2 |
| 2015 | Time Series Classification with Representation Ensembles
Rafael Giusti, Diego Furtado Silva, Gustavo Batista |
IDA | 2 |
| 2015 | Data Stream Classification Guided by Clustering on Nonstationary Environments and Extreme Verification LatencyabstractData stream classification algorithms for nonstationary environments frequently assume the availability of class labels, instantly or with some lag after the classification. However, certain applications, mainly those related to sensors and robotics, involve high costs to obtain new labels during the classification phase. Such a scenario in which the actual labels of processed data are never available is called extreme verification latency. Extreme verification latency requires new classification methods capable of adapting to possible changes over time without external supervision. This paper presents a fast, simple, intuitive and accurate algorithm to classify nonstationary data streams in an extreme verification latency scenario, namely Stream Classification Algorithm Guided by Clustering – SCARGC. Our method consists of a clustering followed by a classification step applied repeatedly in a closed loop fashion. We show in several classification tasks evaluated in synthetic and real data that our method is faster and more accurate than the state-of-the-art. Vinícius M. A. de Souza, Diego Furtado Silva, João Gama 0001, Gustavo Batista |
SDM | 2 |
| 2015 | Class imbalance revisited: a new experimental setup to assess the performance of treatment methods
Ronaldo C. Prati, Gustavo Batista, Diego Furtado Silva |
Knowl. Inf. Syst. | 3 |
| 2014 | Adding Diversity to Rank Examples in Anytime Nearest Neighbor ClassificationabstractIn the last decade we have witnessed a huge increase of interest in data stream learning algorithms. A stream is an ordered sequence of data records. It is characterized by properties such as the potentially infinite and rapid flow of instances. However, a property that is common to various application domains and is frequently disregarded is the very high fluctuating data rates. In domains with fluctuating data rates, the events do not occur with a fixed frequency. This imposes an additional challenge for the classifiers since the next event can occur at any time after the previous one. Anytime classification provides a very convenient approach for fluctuating data rates. In summary, an anytime classifier can be interrupted at any time before its completion and still be able to provide an intermediate solution. The popular k-nearest neighbor (k-NN) classifier can be easily made anytime by introducing a ranking of the training examples. A classification is achieved by scanning the training examples according to this ranking. In this paper, we show how the current state-of-the-art k-NN anytime classifier can be made more accurate by introducing diversity in the training set ranking. Our results show that, with this simple modification, the performance of the anytime version of the k-NN algorithm is consistently improved for a large number of datasets. Cristiano Inácio Lemes, Diego Furtado Silva, Gustavo Batista |
ICMLA | 2 |
| 2014 | Extracting Texture Features for Time Series ClassificationabstractTime series are present in many pattern recognition applications related to medicine, biology, astronomy, economy, and others. In particular, the classification task has attracted much attention from a large number of researchers. In such a task, empirical researches has shown that the 1-Nearest Neighbor rule with a distance measure in time domain usually performs well in a variety of application domains. However, certain time series features are not evident in time domain. A classical example is the classification of sound, in which representative features are usually present in the frequency domain. For these applications, an alternative representation is necessary. In this work we investigate the use of recurrence plots as data representation for time series classification. This representation has well-defined visual texture patterns and their graphical nature exposes hidden patterns and structural changes in data. Therefore, we propose a method capable of extracting texture features from this graphical representation, and use those features to classify time series data. We use traditional methods such as Grey Level Co-occurrence Matrix and Local Binary Patterns, which have shown good results in texture classification. In a comprehensible experimental evaluation, we show that our method outperforms the state-of-the-art methods for time series classification. Vinícius M. A. de Souza, Diego Furtado Silva, Gustavo Batista |
ICPR | 2 |
| 2013 | Time Series Classification Using Compression Distance of Recurrence PlotsabstractThere is a huge increase of interest for time series methods and techniques. Virtually every piece of information collected from human, natural, and biological processes is susceptible to changes over time, and the study of how these changes occur is a central issue in fully understanding such processes. Among all time series mining tasks, classification is likely to be the most prominent one. In time series classification there is a significant body of empirical research that indicates that k-nearest neighbor rule in the time domain is very effective. However, certain time series features are not easily identified in this domain and a change in representation may reveal some significant and unknown features. In this work, we propose the use of recurrence plots as representation domain for time series classification. Our approach measures the similarity between recurrence plots using Campana-Keogh (CK-1) distance, a Kolmogorov complexity-based distance that uses video compression algorithms to estimate image similarity. We show that recurrence plots allied to CK-1 distance lead to significant improvements in accuracy rates compared to Euclidean distance and Dynamic Time Warping in several data sets. Although recurrence plots cannot provide the best accuracy rates for all data sets, we demonstrate that we can predict ahead of time that our method will outperform the time representation with Euclidean and Dynamic Time Warping distances. Diego Furtado Silva, Vinícius M. A. de Souza, Gustavo Batista |
ICDM | 1 |
| 2013 | Applying Machine Learning and Audio Analysis Techniques to Insect Recognition in Intelligent TrapsabstractThroughout the history, insects have had an intimate relationship with humanity, both positive and negative. Insects are vectors of diseases that kill millions of people every year and, at the same time, insects pollinate most of the world's food production. Consequently, there is a demand for new devices able to control the populations of harmful insects while having a minimal impact on beneficial insects. In this paper, we present an intelligent trap that uses a laser sensor to selectively classify and catch insects. We perform an extensive evaluation of different feature sets from audio analysis and machine learning algorithms to construct accurate classifiers for the insect classification task. Support Vector Machines achieved the best results with a MFCC feature set, which consists of coefficients from frequencies scaled according to the human auditory system. We evaluate our classifiers in multiclass and binary class settings, and show that a binary class classifier that recognizes the mosquito species achieved almost perfect accuracy, assuring the applicability of the proposed intelligent trap. Diego Furtado Silva, Vinícius M. A. de Souza, Gustavo Batista, Eamonn J. Keogh, Daniel P. W. Ellis |
ICMLA (1) | 1 |
| 2012 | An Experimental Design to Evaluate Class Imbalance Treatment MethodsabstractIn the last decade, class imbalance has attracted a huge amount of attention from researchers and practitioners. Class imbalance is ubiquitous in Machine Learning, Data Mining and Pattern Recognition applications; therefore, these research communities have responded to such interest with literally dozens of methods and techniques. Surprisingly, there are still many fundamental open-ended questions such as "Are all learning paradigms equally affected by class imbalance?", "What is the expected performance loss for different imbalance degrees?" and "How much of the performance losses can be recovered by the treatment methods?". In this paper, we propose a simple experimental design to assess the performance of class imbalance treatment methods. This experimental setup uses real data sets with artificially modified class distributions to evaluate classifiers in a wide range of class imbalance. We employ such experimental design in a large-scale experimental evaluation with twenty-two data sets and seven learning algorithms from different paradigms. Our results indicate that the expected performance loss, as a percentage of the performance obtained with the balanced distribution, is quite modest (below 5%) for the most balanced distributions up to 10% of minority examples. However, the loss tends to increase quickly for higher degrees of class imbalance, reaching 20% for 1% of minority class examples. Support Vector Machine is the classifier paradigm that is less affected by class imbalance, being almost insensitive to all but the most imbalanced distributions. Finally, we show that the sampling algorithms only partially recover the performance losses. On average, typically about 30\% or less of the performance that was lost due to class imbalance was recovered by random over-sampling and SMOTE. Gustavo Batista, Diego Furtado Silva, Ronaldo C. Prati |
ICMLA (2) | 2 |