VLDB 2026 Research / reviewers in the wild / expert
Jean Paul Barddal
dblp:148/4305
· DBLP profile ↗
56ranked-venue papers
13as first author
30since 2021 · last 2025
0000-0001-9928-854XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 42 · 9 first-author · 24 since 2021Databases, data management, data science and information retrieval · 9 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Adaptive Options for Decision Trees in Evolving Data Stream Classification
Daniel Nowak Assis, Jean Paul Barddal, Fabrício Enembreck |
ECML/PKDD (7) | 2 |
| 2025 | Behavioral insights of adaptive splitting decision trees in evolving data stream classification
Daniel Nowak Assis, Jean Paul Barddal, Fabrício Enembreck |
Knowl. Inf. Syst. | 2 |
| 2025 | Representation ensemble learning applied to facial expression recognition
Bruna Rossetto Delazeri, Andre G. Hochuli, Jean Paul Barddal, Alessandro L. Koerich, Alceu S. Britto Jr. |
Neural Comput. Appl. | 3 |
| 2025 | Concept Drift Adaptation in Text Stream Mining Settings: A Systematic ReviewabstractThe society produces textual data online in several ways, e.g., via reviews and social media posts. Therefore, numerous researchers have been working on discovering patterns in textual data that can indicate peoples’ opinions, interests, and so on. Most tasks regarding natural language processing are addressed using traditional machine learning methods and static datasets. This setting can lead to several problems, e.g., outdated datasets and models, which degrade in performance over time. This is particularly true regarding concept drift, in which the data distribution changes over time. Furthermore, text streaming scenarios also exhibit further challenges, such as the high speed at which data arrive over time. Models for stream scenarios must adhere to the aforementioned constraints while learning from the stream, thus storing texts for limited periods and consuming low memory. This study presents a systematic literature review regarding concept drift adaptation in text stream scenarios. Considering well-defined criteria, we selected 48 papers published between 2018 and August 2024 to unravel aspects such as text drift categories, detection types, model update mechanisms, stream mining tasks addressed, and text representation methods and their update mechanisms. Furthermore, we discussed drift visualization and simulation and listed real-world datasets used in the selected papers. Finally, we brought forward a discussion on existing works in the area, also highlighting open challenges and future research directions for the community. Cristiano Mesquita Garcia, Ramon Abílio, Alessandro L. Koerich, Alceu S. Britto Jr., Jean Paul Barddal |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2024 | LongKey: Keyphrase Extraction for Long DocumentsabstractIn an era of information overload, manually annotating the vast and growing corpus of documents and scholarly papers is increasingly impractical. Automated keyphrase extraction addresses this challenge by identifying representative terms within texts. However, most existing methods focus on short documents (up to 512 tokens), leaving a gap in processing long-context documents. In this paper, we introduce LongKey, a novel framework for extracting keyphrases from lengthy documents, which uses an encoder-based language model to capture extended text intricacies. LongKey uses a max-pooling embedder to enhance keyphrase candidate representation. Validated on the comprehensive LDKP datasets and six diverse, unseen datasets, LongKey consistently outperforms existing unsupervised and language model-based keyphrase extraction methods. Our findings demonstrate LongKey’s versatility and superior performance, marking an advancement in keyphrase extraction for varied text lengths and domains. Jeovane Honório Alves, Radu State, Cinthia Obladen de Almendra Freitas, Jean Paul Barddal |
IEEE Big Data | 4 |
| 2024 | Is it Fine to Tune? Evaluating SentenceBERT Fine-tuning for Brazilian Portuguese Text Stream ClassificationabstractPre-trained language models (LMs) have been used in several scenarios and data mining tasks due to their good-quality representations and their use readiness. Although LMs constitute a significant gain in usability, they are frequently utilized statically over time, meaning that these models can suffer from concept drift and semantic shift, which correspond to changes in data distribution and word meanings. These phenomena are more noticeable when new texts become gradually available. This paper evaluates the impact of updating pre-trained SentenceBERT models overtime on a Brazilian news post classification task in text streaming fashion, a paradigm suitable for learning from data streams. While we update the SBERT model yearly with a reduced number of recent posts, we compare it with scenarios using static LMs. We used the adaptive random forest for classification and evaluated it regarding macro F1-score and elapsed time. The experimental results show that regularly leveraging sampled texts from the recent past for fine-tuning LMs can improve performance metrics over time, reaching better results than using static LMs in most years analyzed. We also evaluated the run times, which suggests that fine-tuning LMs over time provides a good trade-off between performance and run time. Bruno Yuiti Leão Imai, Cristiano Mesquita Garcia, Marcio Vinicius Rocha, Alessandro L. Koerich, Alceu S. Britto Jr., Jean Paul Barddal |
IEEE Big Data | 6 |
| 2024 | Fuels Demand Forecasting: Identifying Leading Feature Sets, Prediction Strategy, and RegressorsabstractFuels are crucial for any country's development and economy, impacting various sectors such as transportation, industry, and electricity generation. Accurate prediction of monthly fuel demand can improve supply chain management, strategic decision-making, and financial planning for businesses while helping governments develop decarbonization policies and estimate pollutant emissions. This paper explores machine learning models to forecast fossil fuels and biofuel demand 12 months ahead, using univariate time series data representing the historical sales of 27 Brazilian states, one of the world's leading producers and consumers of fuels. We evaluate different time series feature sets, machine learning regression models, and prediction strategies to address the complexity of fuel sales influenced by factors such as economic conditions and geopolitical events. Our comprehensive evaluation aims to determine an effective setting for predictive models in the fuel domain. Our results show that popular feature extractors for time series, such as Catch22 and TsFresh, cannot improve the original data representation for most forecasting models. Although focused on Brazil, our findings apply to other countries, since the trained models do not rely on external variables, such as micro and macroeconomic indicators. Jonas Krause, Alexandre C. A. Beiruth, Jean Paul Barddal, Alceu S. Britto Jr., Vinícius M. A. de Souza |
ICMLA | 3 |
| 2024 | Improving Sampling Methods for Fine-Tuning SentenceBERT in Text Streams
Cristiano Mesquita Garcia, Alessandro L. Koerich, Alceu S. Britto Jr., Jean Paul Barddal |
ICPR (19) | 4 |
| 2024 | Alleviating Catastrophic Forgetting in Facial Expression Recognition with Emotion-Centered Models
Israel A. Laurensi R., Alceu S. Britto Jr., Jean Paul Barddal, Alessandro L. Koerich |
ICPR (9) | 3 |
| 2024 | Temporal analysis of drifting hashtags in textual data streams: A graph-based application
Cristiano Mesquita Garcia, Alceu S. Britto Jr., Jean Paul Barddal |
Expert Syst. Appl. | 3 |
| 2023 | Detecting Relevant Information in High- Volume Chat Logs: Keyphrase Extraction for Grooming and Drug Dealing Forensic AnalysisabstractThe growing use of digital communication platforms has given rise to various criminal activities, such as grooming and drug dealing, which pose significant challenges to law en-forcement and forensic experts. This paper presents a supervised keyphrase extraction approach to detect relevant information in high-volume chat logs involving grooming and drug dealing for forensic analysis. The proposed method, JointKPE++, builds upon the JointKPE keyphrase extractor by employing improve-ments to handle longer texts effectively. We evaluate JointKPE++ using BERT-based pre-trained models on grooming and drug dealing datasets, including BERT, RoBERTa, SpanBERT, and BERTimbau. The results show significant improvements over traditional approaches and demonstrate the potential for Join-tKPE++ to aid forensic experts in efficiently detecting keyphrases related to criminal activities. Jeovane Honório Alves, Horácio A. C. G. Pedroso, Rafael Honorio Venetikides, Joel E. M. Köster, Luiz Rodrigo Grochocki, Cinthia Obladen de Almendra Freitas, Jean Paul Barddal |
ICMLA | 7 |
| 2023 | Event-driven Sentiment Drift Analysis in Text Streams: An Application in a Soccer MatchabstractSocial media has been a data source for various applications, given its characteristic of working as a social sensor. Many applications in several areas, such as brand reputation and online opinion monitoring, use this valuable resource to understand the users of services and products. This paper describes an application in the soccer domain, considering data collected from a social media textual data stream. The goal is to detect possible sentiment drifts related to actual events in a soccer match. This task is challenging as we resort to short texts made available during a short time (match length). We evaluated four drift detectors using four metrics: false alarms, delay (considering the number of posts), delay, and missing drifts. Our results show that ADWIN had a stable performance in sentiment drift detection compared to other methods in timely detecting the flagged drifts, raising a small number of false alarms. Given the drifts detected, we used Incremental Word-Vectors to monitor words of interest and check their relatedness to actual events in the match. We empirically assert that the closest words trace back to the sentiment drift generator events. Cristiano Mesquita Garcia, Alceu S. Britto Jr., Jean Paul Barddal |
ICMLA | 3 |
| 2023 | Deep Single Models vs. Ensembles: Insights for a Fast Deployment of Parking Monitoring SystemsabstractSearching for available parking spots in high-density urban centers is a stressful task for drivers that can be mitigated by systems that know in advance the nearest parking space available. To this end, image-based systems offer cost advantages over other sensor-based alternatives (e.g., ultrasonic sensors), requiring less physical infrastructure for installation and maintenance. Despite recent deep learning advances, de-ploying intelligent parking monitoring is still a challenge since most approaches involve collecting and labeling large amounts of data, which is laborious and time-consuming. Our study aims to uncover the challenges in creating a global framework, trained using publicly available labeled parking lot images, that performs accurately across diverse scenarios, enabling the parking space monitoring as a ready-to-use system to deploy in a new environment. Through exhaustive experiments involving different datasets and deep learning architectures, including fusion strategies and ensemble methods, we found that models trained on diverse datasets can achieve 95% accuracy without the burden of data annotation and model training on the target parking lot. Andre G. Hochuli, Jean Paul Barddal, Gillian Cezar Palhano, Leonardo Matheus Mendes, Paulo R. L. Almeida |
ICMLA | 2 |
| 2023 | Mass-Based Short Term Selection of Classifiers in Data StreamsabstractDynamic classifier selection (DCS) regards well-known machine learning techniques in the batch setting that leverage ensemble performance. Most of the methods use similarity-based methods as a proxy, culminating in high computation costs and becoming unfeasible in many streaming scenarios. In this paper, we propose a DCS method able to cope with the high-speed streaming setting, which is based on the performance of base learners in the most recent instances. The impact of our method is evaluated with different ensembles for data streams. We also propose modifications to an Online Boosting method, which has its performance improved with DCS. Our method increases the accuracy and kappa statistic of state-of-the-art ensembles with low overhead of time processing and memory. Daniel Nowak Assis, Fabrício Enembreck, Jean Paul Barddal |
IJCNN | 3 |
| 2023 | Benchmarking Feature Extraction Techniques for Textual Data Stream ClassificationabstractFeature extraction regards transforming unstructured or semi-structured data into structured data that can be used as input for classification and sentiment analysis algorithms, among other applications. This task becomes even more challenging and relevant when textual data becomes available over time as a continuous data stream since the lexicon and semantics can be ever-evolving. Data streams are, by definition, potentially infinite sequences of data that may have ephemeral characteristics, that is, where the data behavior changes, it leads to a phenomenon named concept drift. Textual data streams are specialized data streams, in which texts arrive over time from a continual data source, such as social media, raising challenges in which feature extractors are of great help. In this paper, we benchmark different feature extraction algorithms, i.e., Hashing Trick, Word2Vec, BERT, and Incremental Word-Vectors; in textual data stream classification, considering different stream lengths. The evaluation was performed over a binary and a multiclass classification task, considering two different datasets. Results show that pre-trained models, such as BERT, achieve interesting results, while Hashing Trick also performs competitively. We also observe that incremental methods such as Word2Vec and Incremental Word-Vectors are the most prepared for changing scenarios, yet, they are much more computationally intensive compared to the former when applied to larger streams. Bruno Siedekum Thuma, Pedro Silva de Vargas, Cristiano Mesquita Garcia, Alceu S. Britto Jr., Jean Paul Barddal |
IJCNN | 5 |
| 2023 | An explainable machine learning approach for student dropout prediction
João Gabriel Corrêa Krüger, Alceu S. Britto Jr., Jean Paul Barddal |
Expert Syst. Appl. | 3 |
| 2023 | Incremental specialized and specialized-generalized matrix factorization models based on adaptive learning rate optimizers
Antônio David Viniski, Jean Paul Barddal, Alceu S. Britto Jr., Humberto Vinicius Aparecido de Campos |
Neurocomputing | 2 |
| 2022 | A Machine Learning Approach for School Dropout Prediction in BrazilabstractSchool dropout is a problem that impacts many socioeconomic aspects, including inequality.Dropout prediction algorithms can help remediate this problem, although several past attempts in the literature did so using small datasets.This paper brings forward an experimental approach of machine learning for school dropout prediction in Brazilian schools.The data used for this study was first retrieved from the academic systems of a group of Brazilian private schools, which was later enriched with socio-economic data extracted from governmental sources.Using the dataset to train different types of classifiers, we obtained up to 95.2% precision rates when predicting dropout at different year and educational stages, thus allowing schools to plan and apply retention strategies. João Gabriel Corrêa Krüger, Jean Paul Barddal, Alceu S. Britto Jr. |
ESANN | 2 |
| 2022 | Univariate Time Series Prediction using Data Stream Mining Algorithms and Temporal Dependence
Marcos Alberto Mochinski, Jean Paul Barddal, Fabrício Enembreck |
ICAART (2) | 2 |
| 2022 | Evaluation of Self-taught Learning-based Representations for Facial Emotion RecognitionabstractThis work describes different strategies to generate unsupervised representations obtained through the concept of self-taught learning for facial emotion recognition (FER). The idea is to create complementary representations promoting diver-sity by varying the autoencoders' initialization, architecture, and training data. SVM, Bagging, Random Forest, and a dynamic ensemble selection method are evaluated as final classification methods. Experimental results on JAFFE and Cohn-Kanade datasets using a leave-one-subject-out protocol show that FER methods based on the proposed diverse representations compare favorably against state-of-the-art approaches that also explore unsupervised feature learning. Bruna Rossetto Delazeri, Leonardo León Vera, Jean Paul Barddal, Alessandro L. Koerich, Alceu S. Britto Jr. |
IJCNN | 3 |
| 2022 | Assessing Batch and Online Learning for Delivery in Full and On Time PredictionsabstractImproving results by optimizing process execution is one objective of major companies. For these corporations, the main point for achieving better results is the good maintenance of supply chain management. The most important supply chain metric is Delivery in Full and On Time (DIFOT). DIFOT measures how well a supply chain delivers value to the customer. In this work, we bring forward an analysis of DIFOT prediction from large Brazilian food company. More specifically, we compare a batch and online learning algorithm for DIFOT prediction and depict why the latter is suitable for this problem. Furthermore, we report a feature drift analysis to identify whether there are considerable shifts along with the dataset timespan. As a byproduct of this research, we make the dataset used in this analysis publicly available for future research in DIFOT prediction. Adriano Alves de Lima, Márcie Venâncio Batista, Jean Paul Barddal, Danilo Sipoli Sanches, Luiz Eduardo Soares de Oliveira |
IJCNN | 3 |
| 2022 | Classifying Hierarchical Data Streams using Global Classifiers and Summarization TechniquesabstractThe hierarchical classification of data streams requires models capable of handling a class hierarchy and updating themselves whenever a new example arrives, within restrained processing time and memory consumption. Current state-of-the-art models store raw instances and handle the hierarchy locally, performing a high number of computations at every hierarchy level and with all, eventually redundant, data. This paper introduces Global k-Nearest Centroids (kNC) and Global Dribble, two novel methods for the hierarchical classification of data streams. Both methods use summarization techniques to represent data with constant computational resources usage and a global classification approach to process instances in less time when compared to local strategies. We compare both methods with a state-of-the-art local classifier, and the proposed methods achieved a higher number of correct predictions and process instances nearly twice as fast. Eduardo Tieppo, Jean Paul Barddal, Júlio C. Nievola |
IJCNN | 2 |
| 2022 | Pattern Spotting and Image Retrieval in Historical Documents using Deep HashingabstractThis paper presents a deep learning approach for image retrieval and pattern spotting in digital collections of historical documents. First, a region proposal algorithm detects object candidates in the document page images. Next, deep learning models are used for feature extraction, considering two distinct variants, which provide either real-valued or binary code representations. Finally, candidate images are ranked by computing the feature similarity with a given input query. A robust experimental protocol evaluates the proposed approach considering each representation scheme (real-valued and binary code) on the DocExplore image database. The experimental results show that the proposed deep models compare favorably to the state-of-the-art image retrieval approaches for images of historical documents, outperforming other deep models by 2.56 percentage points using the same techniques for pattern spotting. Besides, the proposed approach also reduces the search time up to 200$\times$, and the storage cost up to 6,000$\times$ when compared to related works based on real-valued representations. Caio da S. Dias, Alceu S. Britto Jr., Jean Paul Barddal, Laurent Heutte, Alessandro L. Koerich |
SMC | 3 |
| 2022 | Improving Data Stream Classification using Incremental Yeo-Johnson Power TransformationabstractData transformation plays an essential role as a preprocessing step in learning models. Several classification techniques have premises about the underlying data distribution, such as normal distribution assumed in Bayesians classifiers. However, applying data transformation in a streaming setting requires processing an infinite and continuous flow of data. In this paper, we propose the Incremental Yeo-Johnson Power Transformation, a variant of the well-known batch Yeo-Johnson transformation that is tailored for streaming settings, i.e., it supports streaming data via statistical sampling and hypothesis testing. Experimental results show that our proposal achieves the same data normality as its batch counterpart. In addition, it improves the prediction performance of a data stream classifier based on Bayesian statistical models. Overall, learning models obtained 3 percentage points improvement. Eduardo Tieppo, Jean Paul Barddal, Júlio C. Nievola |
SMC | 2 |
| 2022 | A systematic review on computer vision-based parking lot management applied on public datasets
Paulo R. L. Almeida, Jeovane Honório Alves, Rafael S. Parpinelli, Jean Paul Barddal |
Expert Syst. Appl. | 4 |
| 2021 | Towards the Overcome of Performance Pitfalls in Data Stream Mining ToolsabstractData stream mining is an essential task in today's scientific community. It allows machine learning models to be updated over time as new data becomes available. Three pillars should be accounted for when selecting an appropriate algorithm for data stream mining: accuracy, processing time, and memory consumption. To develop and assess machine learning models in streaming scenarios, different tools have been developed, where the Massive Online Analysis, written in Java, and scikit-multiflow, written in Python, are in the spotlight. Despite the ease of use of both tools, neither are focused on performance, which puts in jeopardy the usage of the computational resources. In this paper, we show that with the right tools, Python libraries reach performance comparable to C/C++. More specifically, we show how optimized implementations in scikit-multiflow using low-level languages, i.e., C++, C++ with Intel Intrinsics, and Rust; with bindings to Python vastly overcome existing tools in computational resources usage while keeping predictive performance intact. Lucca Portes Cavalheiro, Marco A. Z. Alves, Jean Paul Barddal |
IJCNN | 3 |
| 2021 | Dynamically Selected Ensemble for Data Stream ClassificationabstractMining data streams is a hot topic in the machine learning (ML) community. In addition to learning and updating accurate models over time, these techniques must respect constraints that are not necessarily as strong in batch mode, such as time processing and memory consumption efficiency. A successful family of techniques in batch ML is dynamic classifier selection (DCS). However, these are roughly overlooked in data stream mining. In this paper, we propose a novel dynamic classifier selection framework for data streams called Double Dynamic Classifier Selection (DDCS). We compare DDCS against state-of-art methods for mining data streams in both synthetic and realworld datasets. Results depict that DDCS not only outperforms the state-of-art ensemble methods for data stream classification in terms of accuracy but is also significantly more efficient in terms of processing time and memory consumption. Lucca Portes Cavalheiro, Alceu S. Britto Jr., Jean Paul Barddal, Laurent Heutte |
IJCNN | 3 |
| 2021 | UKIRF: An Item Rejection Framework for Improving Negative Items Sampling in One-Class Collaborative Filtering
Antônio David Viniski, Jean Paul Barddal, Alceu S. Britto Jr. |
PAKDD (2) | 2 |
| 2021 | Adaptive Global k-Nearest Neighbors for Hierarchical Classification of Data StreamsabstractData stream classification differs from batch learning classification methods as data is made available sequentially and may drift over time. Therefore, data stream classification can be simultaneous to all other kinds of classification problems, and it has been revisiting many aspects related to classification in the last years. So far, hierarchical classification was weakly addressed in streaming scenarios despite being a well-established research topic. To fill in this gap between such areas, in this paper, we propose the adaptive global k-Nearest Neighbors for the hierarchical classification of data streams (Global kNN-hDS). Our proposal classifies hierarchical data streams using a constrained memory buffer and a global classification approach. We compare our method against a stateof-the-art local kNN also tailored for streaming scenarios, and results show that our method obtains competitive prediction rates while being statistically faster. Eduardo Tieppo, Jean Paul Barddal, Júlio C. Nievola |
SMC | 2 |
| 2021 | A case study of batch and incremental recommender systems in supermarket data under concept drifts and cold start
Antônio David Viniski, Jean Paul Barddal, Alceu S. Britto Jr., Fabrício Enembreck, Humberto Vinicius Aparecido de Campos |
Expert Syst. Appl. | 2 |
| 2020 | Classifier Pool Generation based on a Two-level Diversity ApproachabstractThis paper describes a classifier pool generation method guided by the diversity estimated on the data complexity and classifier decisions. First, the behavior of complexity measures is assessed by considering several subsamples of the dataset. The complexity measures with high variability across the subsamples are selected for posterior pool adaptation, where an evolutionary algorithm optimizes diversity in both complexity and decision spaces. A robust experimental protocol with 28 datasets and 20 replications is used to evaluate the proposed method. Results show significant accuracy improvements in 69.4% of the experiments when Dynamic Classifier Selection and Dynamic Ensemble Selection methods are applied. Marcos Monteiro 0001, Alceu S. Britto Jr., Jean Paul Barddal, Luiz Eduardo Soares de Oliveira, Robert Sabourin |
ICPR | 3 |
| 2020 | An End-to-End Approach for Recognition of Modern and Historical Handwritten Numeral StringsabstractAn end-to-end solution for handwritten numeral string recognition is proposed, in which the numeral string is considered as composed of objects automatically detected and recognized by a YoLo-based model. The main contribution of this paper is to avoid heuristic-based methods for string preprocessing and segmentation, the need for task-oriented classifiers, and also the use of specific constraints related to the string length. A robust experimental protocol based on several numeral string datasets, including one composed of historical documents, has shown that the proposed method is a feasible end-to-end solution for numeral string recognition. Besides, it reduces the complexity of the string recognition task considerably since it drops out classical steps, in special preprocessing, segmentation, and a set of classifiers devoted to strings with a specific length. Andre G. Hochuli, Alceu S. Britto Jr., Jean Paul Barddal, Robert Sabourin, Luiz Eduardo Soares de Oliveira |
IJCNN | 3 |
| 2020 | Naïve Approaches to Deal With Concept DriftsabstractA common problem in machine learning is to find representative real-world labeled datasets to put the methods to test. When developing approaches to deal with concept drifts, some datasets such as the Forest Covertype and Nebraska Weather are common choices for testing, even though there is no consensus on whether these exhibit concept drifts or not. We argue that some well-known real-world concept drift datasets present a high serial dependence in the target class and may have only minor changes. With this in mind, we propose the use of Naïve methods that should be used for comparison with methods that deal with concept drifts. The experimental results using six real-world well-known concept drift datasets show that the Naïve approaches can be better than some methods to deal with possible concept drifts in datasets such as the Forest Covertype, Electricity, and Nebraska Weather. These results suggest that some widely used datasets may be trivial from the concept drift standpoint, and thus, should be avoided, or at least the results should be compared with the proposed Naïve methods. Paulo R. L. Almeida, Luiz Eduardo Soares de Oliveira, Alceu S. Britto Jr., Jean Paul Barddal |
SMC | 4 |
| 2020 | Combining Slow and Fast Learning for Improved Credit ScoringabstractThe financial credibility of a person is a relevant factor to determine whether a loan should be approved or not, and it is quantified by a credit score, which is computed using past performance on debt obligations, profiling, and other data available. Credit scoring becomes even a hotter topic in emerging countries, as interest rates and customer behavior swiftly vary, given the economic (in)stability of the country and as fintechs are chasing robust solutions for improved credit scoring solutions. Batch machine learning is often deployed for credit scoring, yet, they are tailored for static scenarios, i.e., they are not prepared to swiftly detect and adapt to changes in customer behavior, thus leading to slow recovery in such scenarios. In this paper, we bring forward an analysis on how batch machine learning can be combined with data stream mining techniques, thus leading to better recognition rates in credit scoring scenarios. We analyze three different real-world datasets from Brazilian financial institutions, whilst keeping their secrecy preserved, and show how batch and stream learning can be combined towards improved credit scoring systems, as well as highlighting relevant gaps that still require attention. Jean Paul Barddal, Fabrício Enembreck, Lucas Loezer, Riccardo Lanzuolo |
SMC | 1 |
| 2020 | ADADRIFT: An Adaptive Learning Technique for Long-history Stream-based Recommender SystemsabstractAdaptive recommender systems are increasingly showing their importance as profiling is a dynamic problem. Their goal is to update recommendation models as new interactions take place, thus swiftly adapting to drifts in the user's behavior and desires, and item's audience. However, existing recommendation algorithms usually do not perform well during drifts, as they take long to adapt to changes, or these updates are suboptimal since they account for all profiles' preferences equally, which is often untrue as each individual and its changes are unique. In this paper, we propose the ADADRIFT algorithm to deal with user and item-based drifts in adaptive recommender systems using personalized learning rates based on profile statistics. The experiments using stream-based recommender systems (ISGD and BRISMF) across four different datasets show that ADADRIFT surpasses ADADELTA with significant improvements in recommendation rates. The best results appear when the data streams have a long history of the users' or items' interactions and drifts become noticeable. The experimentation in this work highlight the importance of handling drifts in recommender systems. Eduardo Ferreira José, Fabrício Enembreck, Jean Paul Barddal |
SMC | 3 |
| 2020 | Improving Multiple Time Series Forecasting with Data Stream Mining AlgorithmsabstractThis paper proposes a hybrid ensemble learning approach that combines statistical and data stream mining algorithms to obtain better forecasting performance in multiple time series prediction problems. Although some multiple time series algorithms perform surprisingly well in a variety of domains, it is well-known that no one is dominant for every existent domain. Therefore, we developed a meta-technique based on data stream mining and static ensemble selection strategy and evaluated its forecasting goodness-of-fit in time series datasets from M3 and M4 competitions. After training different regression models, we show how the combination of auto.arima and AdaGrad leads to improved forecasting rates, thus surpassing the results of state-of-art algorithms. Marcos Alberto Mochinski, Jean Paul Barddal, Fabrício Enembreck |
SMC | 2 |
| 2020 | Lessons learned from data stream classification applied to credit scoring
Jean Paul Barddal, Lucas Loezer, Fabrício Enembreck, Riccardo Lanzuolo |
Expert Syst. Appl. | 1 |
| 2019 | Vertical and Horizontal Partitioning in Data Stream Regression EnsemblesabstractData stream mining is an emerging topic in machine learning that targets the creation and update of predictive models over time as new data becomes available. Regarding existing works, classification is the most widely tackled task, which leaves regression nearly untouched. In this paper, the focus relies on ensemble learning for data stream regression, more specifically on vertical and horizontal data partitioning techniques. The goal is to determine whether and under which conditions partitioning can lessen the error rates of different types of learners in the data stream regression task. The proposed method combines vertical and horizontal partitioning, and it is compared with and against different types of learners and existing ensembles. Jean Paul Barddal |
IJCNN | 1 |
| 2019 | Merit-guided dynamic feature selection filter for data streams
Jean Paul Barddal, Fabrício Enembreck, Heitor Murilo Gomes, Albert Bifet, Bernhard Pfahringer |
Expert Syst. Appl. | 1 |
| 2019 | Boosting decision stumps for dynamic feature selection on data streams
Jean Paul Barddal, Fabrício Enembreck, Heitor Murilo Gomes, Albert Bifet, Bernhard Pfahringer |
Inf. Syst. | 1 |
| 2019 | Correction to: Adaptive random forests for evolving data stream classification
Heitor Murilo Gomes, Albert Bifet, Jesse Read, Jean Paul Barddal, Fabrício Enembreck, Bernhard Pfahringer, Geoff Holmes 0001, Talel Abdessalem |
Mach. Learn. | 4 |
| 2018 | Adaptive random forests for data stream regression
Heitor Murilo Gomes, Jean Paul Barddal, Luis Eduardo Boiko Ferreira, Albert Bifet |
ESANN | 2 |
| 2018 | An Experimental Perspective on Sampling Methods for Imbalanced Learning From Financial DatabasesabstractThe financial market is one of the major consumers of data mining techniques, and the main reason is their efficiency to analyze complex data. One important trait shared between most financial applications is class imbalance. Since traditional classification methods assume nearly balanced classes and equal misclassification costs, they usually fail to deal with imbalanced data. However, in financial contexts, problems are usually imbalanced, and instances from the minority class are known for deficits of millions of dollars every year, e.g., credit card frauds, money laundering transactions and so forth. Over the years, several techniques for dealing with class imbalance have been developed, such as sampling techniques and algorithm adaptations. In this study, we analyze how different sampling techniques impact the performance of different classification systems on financial applications. Results show that, for the given datasets, sampling techniques allow the improvement of prediction performance of the minority class while also improving overall classification rates. Nevertheless, their use often deteriorates the performance in predicting the majority class. Luis Eduardo Boiko Ferreira, Jean Paul Barddal, Fabrício Enembreck, Heitor Murilo Gomes |
IJCNN | 2 |
| 2018 | Are fintechs really a hype? A machine learning-based polarity analysis of Brazilian posts on social mediaabstractFintechs are technology companies that, in contrast to traditional banks, are engaged in digital solutions for payment, money transfers, and real-time notifications. Taking advantage of digital means of communication, most of the service interactions between fintechs and customers occurs via chats or posts in social media. In this work, our goal is to use machine learning to analyze these posts and identify what are the terms used by customers to express positive, neutral and negative customer experiences. During this analysis, we assess the following questions using data from the 3 biggest fintechs in Brazil: (i) what are the most commented topics on social media regarding fintechs, (ii) what are the words more often used by customers to express positive, negative and neutral reactions to the customer service obtained; and (iii) what kind of machine learning model should a fintech use to automatically identify whether a post is positive, negative or neutral. Marina Ponestke Seara, Andreia Malucelli, Altair Olivo Santin, Jean Paul Barddal |
INDIN | 4 |
| 2017 | Improving Credit Risk Prediction in Online Peer-to-Peer (P2P) Lending Using Imbalanced Learning TechniquesabstractPeer-to-peer (P2P) lending is a global trend of financial markets that allow individuals to obtain and concede loans without having financial institutions as a strong proxy. As many real-world applications, P2P lending presents an imbalanced characteristic, where the number of creditworthy loan requests is much larger than the number of non-creditworthy ones. In this work, we wrangle a real-world P2P lending data set from Lending Club, containing a large amount of data gathered from 2007 up to 2016. We analyze how supervised classification models and techniques to handle class imbalance impact creditworthiness prediction rates. Ensembles, cost-sensitive and sampling methods are combined and evaluated along logistic regression, decision tree, and bayesian learning schemes. Results show that, in average, sampling techniques outperform ensembles and cost sensitive approaches. Luis Eduardo Boiko Ferreira, Jean Paul Barddal, Heitor Murilo Gomes, Fabrício Enembreck |
ICTAI | 2 |
| 2017 | A survey on feature drift adaptation: Definition, benchmark, challenges and future directions
Jean Paul Barddal, Heitor Murilo Gomes, Fabrício Enembreck, Bernhard Pfahringer |
J. Syst. Softw. | 1 |
| 2017 | Adaptive random forests for evolving data stream classification
Heitor Murilo Gomes, Albert Bifet, Jesse Read, Jean Paul Barddal, Fabrício Enembreck, Bernhard Pfahringer, Geoff Holmes 0001, Talel Abdessalem |
Mach. Learn. | 4 |
| 2016 | A benchmark of classifiers on feature drifting data streamsabstractThe ever increasing data generation confronts both practitioners and researchers on handling massive and sequentially generated amounts of information, the so-called data streams. In this context, a lot of effort has been put on the extraction of useful patterns from streaming scenarios. Learning from data streams embeds a variety of problems, and by far, the most challenging is concept drift, i.e. changes in data distribution. In this paper, we focus on a specific type of drift uncommonly assessed in the literature: feature drifts. Feature drifts occur whenever a subset of features becomes, or ceases to be, relevant to the concept to be learned. We propose and review several feature drifting data stream generators and use them to benchmark state-of-the-art data stream classification algorithms and their combination with drift detectors. Results show that, although drift detectors enable slight quicker recovery to feature drifts, best results are obtained by Hoeffding Adaptive Tree, the only learner that performs dynamic feature selection as streams progress. Jean Paul Barddal, Heitor Murilo Gomes, Alceu S. Britto Jr., Fabrício Enembreck |
ICPR | 1 |
| 2016 | Overcoming feature drifts via dynamic feature weighted k-nearest neighbor learningabstractExtracting useful knowledge from data streams is problematic, mainly due to changes in their data distribution, a phenomenon named concept drift. Recently, studies have shown that most of existing algorithms for learning from data streams do not encompass techniques for a specific kind of drift: feature drifts. Feature drifts occur when features become, or cease to be, relevant to the learning task. In this paper, we propose an extension to the k-nearest neighbor classifier, so its distances' computations are weighted according to their current discriminative power. On our proposal, the discriminative power of features is given by entropy, which is swiftly computed over a sliding window. Empirical evidence shows that our approach is able to overcome several existing algorithms in accuracy and feature drift adaptation, while at the expense of bounded processing time and memory space. Jean Paul Barddal, Heitor Murilo Gomes, Jones Granatyr, Alceu S. Britto Jr., Fabrício Enembreck |
ICPR | 1 |
| 2016 | Towards emotion-based reputation guessing learning agentsabstractTrust and reputation mechanisms are part of the logical protection of intelligent agents, preventing malicious agents from acting egotistically or with the intention to damage others. Several studies in Psychology, Neurology and Anthropology claim that emotions are part of human's decision making process. However, there is a lack of understanding about how affective aspects, such as emotions, influence trust or reputation levels of intelligent agents when they are inserted into an information exchange environment, e.g. an evaluation system. In this paper we propose a reputation model that accounts for emotional bounds given by Ekman's basic emotions and inductive machine learning. Our proposal is evaluated by extracting emotions from texts provided by two online human-fed evaluation systems. Empirical results show significant agent's utility improvements with p <; .05 when compared to non-emotion-wise proposals, thus, showing the need for future research in this area. Jones Granatyr, Jean Paul Barddal, Adriano Weihmayer Almeida, Fabrício Enembreck, Adaiane Pereira dos Santos Granatyr |
IJCNN | 2 |
| 2016 | On Dynamic Feature Weighting for Feature Drifting Data Streams
Jean Paul Barddal, Heitor Murilo Gomes, Fabrício Enembreck, Bernhard Pfahringer, Albert Bifet |
ECML/PKDD (2) | 1 |
| 2016 | SNCStream+: Extending a high quality true anytime data stream clustering algorithm
Jean Paul Barddal, Heitor Murilo Gomes, Fabrício Enembreck, Jean-Paul A. Barthès |
Inf. Syst. | 1 |
| 2015 | Analyzing the Impact of Feature Drifts in Streaming Learning
Jean Paul Barddal, Heitor Murilo Gomes, Fabrício Enembreck |
ICONIP (1) | 1 |
| 2015 | A Complex Network-Based Anytime Data Stream Clustering Algorithm
Jean Paul Barddal, Heitor Murilo Gomes, Fabrício Enembreck |
ICONIP (1) | 1 |
| 2015 | On the Discovery of Time Distance Constrained Temporal Association Rules
Heitor Murilo Gomes, Deborah Ribeiro de Carvalho, Lourdes Zubieta, Jean Paul Barddal, Andreia Malucelli |
ICONIP (2) | 4 |
| 2015 | A Survey on Feature Drift AdaptationabstractMining data streams is of the utmost importance due to its appearance in many real-world situations, such as: sensor networks, stock market analysis and computer networks intrusion detection systems. Data streams are, by definition, potentially unbounded sequences of data that arrive intermittently at rapid rates. Extracting useful knowledge from data streams embeds virtually all problems from conventional data mining with the addition of single-pass real-time processing within limited time and memory space. Additionally, due to its ephemeral nature, it is expected that streams undergo changes in its data distribution denominated concept drifts. In this work, we focus on one specific kind of concept drift that has not been extensively addressed in the literature, namely feature drift. A feature drift happens when changes occur in the set of features, such that a subset of features become, or cease to be, relevant to the learning problem. Specifically, changes in the relevance of features directly imply modifications in the decision boundary to be learned, thus the learner must detect and adapt to according to it. Timely detection and recover from feature drifts is a challenging task that can be modeled after a dynamic feature selection problem. In this paper we survey existing work on dynamic feature selection for data streams that acts either implicitly or explicitly. We conclude that there is a need for future research in this area, which we highlight as future research directions. Jean Paul Barddal, Heitor Murilo Gomes, Fabrício Enembreck |
ICTAI | 1 |