EDBT 2026 Demo / reviewers in the wild / expert
Piotr Duda
dblp:94/11189
· DBLP profile ↗
9ranked-venue papers in the field
2as first author
3since 2021 · last 2027
0000-0001-7182-1349ORCID · conflict
Domains — venue-derived; a paper can count in several
Knowledge Engineering, Semantic Web & Information Systems · 7 (2 first)Database Systems & Data Management · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Towards faster and deeper hoeffding trees: Fractional bounds and correlation-preserving multi-index statisticsabstractThis paper presents two complementary enhancements to the Very Fast Decision Tree algorithm for data stream mining. The first contribution introduces the fractional Hoeffding bound, a relaxed splitting criterion where the original threshold is scaled by a factor . Experimental evidence shows that for this modification, the resulting trees not only grow faster but also achieve higher accuracy compared to the original Very Fast Decision tree. The second contribution proposes a novel data structure, called extended statistics, which extends the traditional sufficient statistics used in the Very Fast Decision Tree by maintaining additional information about attribute co-occurrences. This allows child nodes to inherit richer knowledge from their parent, leading to deeper trees with accelerated convergence to the target accuracy. Numerical experiments on synthetic data streams demonstrate that the combination of fractional bounds and multi-index statistics yields significant accuracy gains, particularly in the early phases of learning. The extended statistics improvement, however, comes at the cost of increased memory and computational requirements, emphasizing the trade-off between predictive performance and resource usage. Maciej Jaworski, Danuta Rutkowska, Piotr Duda, Xinyu Geng, Leszek Rutkowski |
Inf. Sci. | 3 |
| 2024 | Accelerating deep neural network learning using data stream methodology
Piotr Duda, Mateusz Wojtulewicz, Leszek Rutkowski |
Inf. Sci. | 1 |
| 2023 | The L2 convergence of stream data mining algorithms based on probabilistic neural networks
Danuta Rutkowska, Piotr Duda, Jinde Cao, Leszek Rutkowski, Aleksander Byrski, Maciej Jaworski, Dacheng Tao |
Inf. Sci. | 2 |
| 2019 | Corrigendum to 'How to adjust an ensemble size in stream data mining?' Information Sciences, vol. 381 (2017), pp. 46-54
Lena Pietruczuk, Leszek Rutkowski, Maciej Jaworski, Piotr Duda |
Inf. Sci. | 4 |
| 2018 | Knowledge discovery in data streams with the orthogonal series-based generalized regression neural networks
Piotr Duda, Maciej Jaworski, Leszek Rutkowski |
Inf. Sci. | 1 |
| 2017 | How to adjust an ensemble size in stream data mining?
Lena Pietruczuk, Leszek Rutkowski, Maciej Jaworski, Piotr Duda |
Inf. Sci. | 4 |
| 2014 | The CART decision tree for mining data streams
Leszek Rutkowski, Maciej Jaworski, Lena Pietruczuk, Piotr Duda |
Inf. Sci. | 4 |
| 2014 | Decision Trees for Mining Data Streams Based on the Gaussian ApproximationabstractSince the Hoeffding tree algorithm was proposed in the literature, decision trees became one of the most popular tools for mining data streams. The key point of constructing the decision tree is to determine the best attribute to split the considered node. Several methods to solve this problem were presented so far. However, they are either wrongly mathematically justified (e.g., in the Hoeffding tree algorithm) or time-consuming (e.g., in the McDiarmid tree algorithm). In this paper, we propose a new method which significantly outperforms the McDiarmid tree algorithm and has a solid mathematical basis. Our method ensures, with a high probability set by the user, that the best attribute chosen in the considered node using a finite data sample is the same as it would be in the case of the whole data stream. Leszek Rutkowski, Maciej Jaworski, Lena Pietruczuk, Piotr Duda |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2013 | Decision Trees for Mining Data Streams Based on the McDiarmid's BoundabstractIn mining data streams the most popular tool is the Hoeffding tree algorithm. It uses the Hoeffding's bound to determine the smallest number of examples needed at a node to select a splitting attribute. In the literature the same Hoeffding's bound was used for any evaluation function (heuristic measure), e.g., information gain or Gini index. In this paper, it is shown that the Hoeffding's inequality is not appropriate to solve the underlying problem. We prove two theorems presenting the McDiarmid's bound for both the information gain, used in ID3 algorithm, and for Gini index, used in Classification and Regression Trees (CART) algorithm. The results of the paper guarantee that a decision tree learning system, applied to data streams and based on the McDiarmid's bound, has the property that its output is nearly identical to that of a conventional learner. The results of the paper have a great impact on the state of the art of mining data streams and various developed so far methods and algorithms should be reconsidered. Leszek Rutkowski, Lena Pietruczuk, Piotr Duda, Maciej Jaworski |
IEEE Trans. Knowl. Data Eng. | 3 |