EDBT 2026 Demo / reviewers in the wild / expert
Yibin Sun
dblp:260/1305
· DBLP profile ↗
7ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0002-8325-1889ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 5 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive approaches towards fully incremental prediction interval for data stream regressionabstractAbstract Prediction intervals (PIs) are a practical tool for uncertainty quantification in regression, but comparatively little work has addressed fully incremental PI generation for data streams. In streaming settings, data arrive continuously, each instance is typically processed once, and concept drift can quickly invalidate a previously well-calibrated interval. These properties make many batch PI methods and window-based adaptations difficult to apply efficiently. This paper studies Adaptive Prediction Interval (AdaPI), an online post-calibration framework that adjusts interval width according to observed coverage. We instantiate the framework with a fully incremental variant of Mean and Variance Estimation (MVE) and investigate three adaptive scaling functions. We also adopt an evaluation perspective that jointly considers coverage accuracy and interval width. Experiments on a collection of real-world and synthetic regression streams show that AdaPI can often move coverage closer to the desired confidence level while maintaining competitive interval width; under the default 95% confidence setting and coverage-heavy CING weighting, the linear variant frequently gives the strongest empirical coverage–width trade-off among the three adaptive strategies. Yibin Sun, Bernhard Pfahringer, Heitor Murilo Gomes, Albert Bifet |
Knowl. Inf. Syst. | 1 |
| 2025 | Dynamic Ensemble Member Selection for Data Stream ClassificationabstractEnsemble methods are widely recognized for their effectiveness in data stream classification. This paper introduces Dynamic Ensemble Member Selection (DEMS), a novel framework that dynamically selects a subset of classifiers from an ensemble for each individual prediction. DEMS ranks base learners based on estimated accuracy and predictive margin, using only the top-K members for prediction, where K is optimized in a self-adaptive manner. The proposed method significantly enhances predictive performance across various state-of-the-art ensemble algorithms, including Streaming Random Patches, Adaptive Random Forest, and Online Smooth Boost. Experimental results demonstrate that DEMS consistently improves classification accuracy while maintaining a minimal runtime overhead of just 11.66% compared to the original methods. This work highlights the potential of DEMS in adapting to concept drift and optimizing ensemble diversity, offering a practical solution for real-time data stream classification. Yibin Sun, Bernhard Pfahringer, Heitor Murilo Gomes, Albert Bifet |
CIKM | 1 |
| 2025 | Machine Learning on the Fly: A Hands-On Tutorial for Streaming DataabstractData stream learning is an emerging machine learning paradigm designed for environments where data arrive continuously and must be processed in real time. Unlike traditional batch learning, which assumes access to a fixed dataset, stream learning addresses the unique challenges of non-stationary distributions, bounded memory, and strict computational constraints. These challenges are increasingly relevant across domains such as IoT, finance, cybersecurity, and environmental monitoring, where timely and adaptive decision-making is essential. This tutorial introduces key concepts and techniques in data stream learning, blending foundational theory with practical demonstrations. It features CapyMOA, an open-source library that provides efficient algorithm implementations through a high-level Python API. We demonstrate the use of this tool through practical examples, with all source code available at https://github.com/adaptive-machine-learning/CapyMOA, and supporting tutorials and installation guides accessible at https://capymoa.org/. Heitor Murilo Gomes, Nuwan Gunasekara, Yibin Sun |
ICDE | 3 |
| 2025 | Detecting Domain Shifts in Myoelectric Activations: Challenges and Opportunities in Stream Learning
Yibin Sun, Nick Jin Sean Lim, Guilherme Weigert Cassales, Heitor Murilo Gomes, Bernhard Pfahringer, Albert Bifet, Anany Dwivedi |
PRICAI (5) | 1 |
| 2024 | Adaptive Prediction Interval for Data Stream Regression
Yibin Sun, Bernhard Pfahringer, Heitor Murilo Gomes, Albert Bifet |
PAKDD (3) | 1 |
| 2024 | Real-Time Energy Pricing in New Zealand: An Evolving Stream Analysis
Yibin Sun, Heitor Murilo Gomes, Bernhard Pfahringer, Albert Bifet |
PRICAI (5) | 1 |
| 2022 | SOKNL: A novel way of integrating K-nearest neighbours with adaptive random forest regression for data streamsabstractAbstract Most research in machine learning for data streams has focused on classification algorithms, whereas regression methods have received a lot less attention. This paper proposes Self-Optimising K-Nearest Leaves (SOKNL), a novel forest-based algorithm for streaming regression problems. Specifically, the Adaptive Random Forest Regression, a state-of-the-art online regression algorithm is extended like this: in each leaf, a representative data point – also called centroid – is generated by compressing the information from all instances in that leaf. During the prediction step, instead of letting all trees in the forest participate, the distances between the input instance and all centroids from relevant leaves are calculated, only k trees that possess the smallest distances are utilised for the prediction. Furthermore, we simplify the algorithm by introducing a mechanism for tuning the k values, which is dynamically and automatically optimised based on historical information. This new algorithm produces promising predictive results and achieves a superior ranking according to statistical testing when compared with several standard stream regression methods over typical benchmark datasets. This improvement incurs only a small increase in runtime and memory consumption over the basic Adaptive Random Forest Regressor. Yibin Sun, Bernhard Pfahringer, Heitor Murilo Gomes, Albert Bifet |
Data Min. Knowl. Discov. | 1 |