EDBT 2026 Demo / reviewers in the wild / expert
Jean Paul Barddal
dblp:148/4305
· DBLP profile ↗
9ranked-venue papers in the field
3as first author
6since 2021 · last 2025
0000-0001-9928-854XORCID · verified
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 5 (1 first)Database Systems & Data Management · 2 (2 first)Big Data, Cloud & Distributed Data Systems · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Adaptive Options for Decision Trees in Evolving Data Stream Classification
Daniel Nowak Assis, Jean Paul Barddal, Fabrício Enembreck |
ECML/PKDD (7) | 2 |
| 2025 | Behavioral insights of adaptive splitting decision trees in evolving data stream classification
Daniel Nowak Assis, Jean Paul Barddal, Fabrício Enembreck |
Knowl. Inf. Syst. | 2 |
| 2025 | Concept Drift Adaptation in Text Stream Mining Settings: A Systematic ReviewabstractThe society produces textual data online in several ways, e.g., via reviews and social media posts. Therefore, numerous researchers have been working on discovering patterns in textual data that can indicate peoples’ opinions, interests, and so on. Most tasks regarding natural language processing are addressed using traditional machine learning methods and static datasets. This setting can lead to several problems, e.g., outdated datasets and models, which degrade in performance over time. This is particularly true regarding concept drift, in which the data distribution changes over time. Furthermore, text streaming scenarios also exhibit further challenges, such as the high speed at which data arrive over time. Models for stream scenarios must adhere to the aforementioned constraints while learning from the stream, thus storing texts for limited periods and consuming low memory. This study presents a systematic literature review regarding concept drift adaptation in text stream scenarios. Considering well-defined criteria, we selected 48 papers published between 2018 and August 2024 to unravel aspects such as text drift categories, detection types, model update mechanisms, stream mining tasks addressed, and text representation methods and their update mechanisms. Furthermore, we discussed drift visualization and simulation and listed real-world datasets used in the selected papers. Finally, we brought forward a discussion on existing works in the area, also highlighting open challenges and future research directions for the community. Cristiano Mesquita Garcia, Ramon Abílio, Alessandro L. Koerich, Alceu S. Britto Jr., Jean Paul Barddal |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2024 | LongKey: Keyphrase Extraction for Long DocumentsabstractIn an era of information overload, manually annotating the vast and growing corpus of documents and scholarly papers is increasingly impractical. Automated keyphrase extraction addresses this challenge by identifying representative terms within texts. However, most existing methods focus on short documents (up to 512 tokens), leaving a gap in processing long-context documents. In this paper, we introduce LongKey, a novel framework for extracting keyphrases from lengthy documents, which uses an encoder-based language model to capture extended text intricacies. LongKey uses a max-pooling embedder to enhance keyphrase candidate representation. Validated on the comprehensive LDKP datasets and six diverse, unseen datasets, LongKey consistently outperforms existing unsupervised and language model-based keyphrase extraction methods. Our findings demonstrate LongKey’s versatility and superior performance, marking an advancement in keyphrase extraction for varied text lengths and domains. Jeovane Honório Alves, Radu State, Cinthia Obladen de Almendra Freitas, Jean Paul Barddal |
IEEE Big Data | 4 |
| 2024 | Is it Fine to Tune? Evaluating SentenceBERT Fine-tuning for Brazilian Portuguese Text Stream ClassificationabstractPre-trained language models (LMs) have been used in several scenarios and data mining tasks due to their good-quality representations and their use readiness. Although LMs constitute a significant gain in usability, they are frequently utilized statically over time, meaning that these models can suffer from concept drift and semantic shift, which correspond to changes in data distribution and word meanings. These phenomena are more noticeable when new texts become gradually available. This paper evaluates the impact of updating pre-trained SentenceBERT models overtime on a Brazilian news post classification task in text streaming fashion, a paradigm suitable for learning from data streams. While we update the SBERT model yearly with a reduced number of recent posts, we compare it with scenarios using static LMs. We used the adaptive random forest for classification and evaluated it regarding macro F1-score and elapsed time. The experimental results show that regularly leveraging sampled texts from the recent past for fine-tuning LMs can improve performance metrics over time, reaching better results than using static LMs in most years analyzed. We also evaluated the run times, which suggests that fine-tuning LMs over time provides a good trade-off between performance and run time. Bruno Yuiti Leão Imai, Cristiano Mesquita Garcia, Marcio Vinicius Rocha, Alessandro L. Koerich, Alceu S. Britto Jr., Jean Paul Barddal |
IEEE Big Data | 6 |
| 2021 | UKIRF: An Item Rejection Framework for Improving Negative Items Sampling in One-Class Collaborative Filtering
Antônio David Viniski, Jean Paul Barddal, Alceu S. Britto Jr. |
PAKDD (2) | 2 |
| 2019 | Boosting decision stumps for dynamic feature selection on data streams
Jean Paul Barddal, Fabrício Enembreck, Heitor Murilo Gomes, Albert Bifet, Bernhard Pfahringer |
Inf. Syst. | 1 |
| 2016 | On Dynamic Feature Weighting for Feature Drifting Data Streams
Jean Paul Barddal, Heitor Murilo Gomes, Fabrício Enembreck, Bernhard Pfahringer, Albert Bifet |
ECML/PKDD (2) | 1 |
| 2016 | SNCStream+: Extending a high quality true anytime data stream clustering algorithm
Jean Paul Barddal, Heitor Murilo Gomes, Fabrício Enembreck, Jean-Paul A. Barthès |
Inf. Syst. | 1 |