Maciej Grzenda

dblp:01/1747 · DBLP profile ↗
← Back
5ranked-venue papers in the field
3as first author
4since 2021 · last 2026
0000-0002-5440-4954ORCID · corroborated

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 4 (3 first)Database Systems & Data Management · 1
YearPublicationVenuePosition
2026 SLEADE: Disagreement-Based Semi-Supervised Learning for Sparsely Labeled Evolving Data Streams
abstract
Semi-supervised learning (SSL) problems are challenging, appear in many domains, and are particularly relevant to streaming applications, where data are abundant but labels are not. The problem tackled here is classification over an evolving data stream where labels are rare and distributed randomly. We propose SLEADE (Stream LEArning by Disagreement Ensemble), a novel method that exploits disagreement-based learning and unsupervised drift detection to leverage unlabeled data during training. SLEADE uses pseudo-labeled instances to augment the training set of each member of an ensemble using amajority trains minorityscheme. The pseudo-labeled data impact is controlled by a weighting function that considers the confidence in the prediction attributed by the ensemble members. SLEADE exploits unsupervised drift detection, which allows the ensemble to respond to changes. We present several experiments using real and synthetic data to illustrate the benefits and limitations of SLEADE compared to existing algorithms.
Heitor Murilo Gomes, Jesse Read, Maciej Grzenda, Bernhard Pfahringer, Albert Bifet
IEEE Trans. Knowl. Data Eng.3
2024 Hybrid Ensemble-Based Travel Mode Prediction
Pawel Golik, Maciej Grzenda, Elzbieta Sienkiewicz
IDA (1)2
2022 Quantifying Changes in Predictions of Classification Models for Data Streams
Maciej Grzenda
IDA1
2022 Urban Traveller Preference Miner: Modelling Transport Choices with Survey Data Streams
Maciej Grzenda, Marcin Luckner, Przemyslaw Wrona
ECML/PKDD (6)1
2020 Delayed labelling evaluation for data streams
abstract
Abstract A large portion of the stream mining studies on classification rely on the availability of true labels immediately after making predictions. This approach is well exemplified by the test-then-train evaluation, where predictions immediately precede true label arrival. However, in many real scenarios, labels arrive with non-negligible latency. This raises the question of how to evaluate classifiers trained in such circumstances. This question is of particular importance when stream mining models are expected to refine their predictions between acquiring instance data and receiving its true label. In this work, we propose a novel evaluation methodology for data streams when verification latency takes place, namely continuous re-evaluation. It is applied to reference data streams and it is used to differentiate between stream mining techniques in terms of their ability to refine predictions based on newly arriving instances. Our study points out, discusses and shows empirically the importance of considering the delay of instance labels when evaluating classifiers for data streams.
Maciej Grzenda, Heitor Murilo Gomes, Albert Bifet
Data Min. Knowl. Discov.1