VLDB 2026 Research / reviewers in the wild / expert
Piotr Duda
dblp:94/11189
· DBLP profile ↗
21ranked-venue papers
6as first author
4since 2021 · last 2027
0000-0001-7182-1349ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 first-author · 1 since 2021Databases, data management, data science and information retrieval · 9 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Towards faster and deeper hoeffding trees: Fractional bounds and correlation-preserving multi-index statisticsabstractThis paper presents two complementary enhancements to the Very Fast Decision Tree algorithm for data stream mining. The first contribution introduces the fractional Hoeffding bound, a relaxed splitting criterion where the original threshold is scaled by a factor . Experimental evidence shows that for this modification, the resulting trees not only grow faster but also achieve higher accuracy compared to the original Very Fast Decision tree. The second contribution proposes a novel data structure, called extended statistics, which extends the traditional sufficient statistics used in the Very Fast Decision Tree by maintaining additional information about attribute co-occurrences. This allows child nodes to inherit richer knowledge from their parent, leading to deeper trees with accelerated convergence to the target accuracy. Numerical experiments on synthetic data streams demonstrate that the combination of fractional bounds and multi-index statistics yields significant accuracy gains, particularly in the early phases of learning. The extended statistics improvement, however, comes at the cost of increased memory and computational requirements, emphasizing the trade-off between predictive performance and resource usage. Maciej Jaworski, Danuta Rutkowska, Piotr Duda, Xinyu Geng, Leszek Rutkowski |
Inf. Sci. | 3 |
| 2026 | Improving noisy label learning via transition matrix estimation from feature embeddings
Mateusz Wojtulewicz, Piotr Duda, Dacheng Tao, Leszek Rutkowski |
Knowl. Based Syst. | 2 |
| 2024 | Accelerating deep neural network learning using data stream methodology
Piotr Duda, Mateusz Wojtulewicz, Leszek Rutkowski |
Inf. Sci. | 1 |
| 2023 | The L2 convergence of stream data mining algorithms based on probabilistic neural networks
Danuta Rutkowska, Piotr Duda, Jinde Cao, Leszek Rutkowski, Aleksander Byrski, Maciej Jaworski, Dacheng Tao |
Inf. Sci. | 2 |
| 2020 | Analysis of microtomographic images in automatic defect localization and detectionabstractAbstract The paper presents a fast method of fully automatic localization and classification of defects in aluminium castings based on computed microtomography images. In the light of current research and based on available publications, where such analysis is made on the basis of images obtained from standard radiography (x-ray), this is a new approach which uses microtomographic images ( $$\mu $$ μ -CT). In addition, the above-mentioned solutions most often analyze a pre-separated portion of an image, which requires the initial operator interference. The authors’ own pre-processing methods, which allow to separate the element area and potential defect areas from $$\mu $$ μ -CT images, and methods of extraction of selected features describing these areas have been proposed in the solution discussed here. A neural network trained using the Levenberg–Marquardt method with error backpropagation has been used as a classifier. The optimal network structure 20–4–1 and a set of 20 features describing the analysed areas have been determined as a result of performed tests. The applied solutions have provided 89% correct detection for any defect size and 96.73% for large defects, which is comparable to the results obtained from methods using x-ray images. This has confirmed that it is possible to use $$\mu $$ μ -CT images in automatic defect localization in 3D. Thanks to this method, quantitative analysis of aluminium castings can be carried out without user interaction and fully automated. Mariusz Marzec, Piotr Duda, Zygmunt Wróbel |
Mach. Vis. Appl. | 2 |
| 2020 | On the Parzen Kernel-Based Probability Density Function Learning Procedures Over Time-Varying Streaming Data With Applications to Pattern ClassificationabstractIn this paper, we propose a recursive variant of the Parzen kernel density estimator (KDE) to track changes of dynamic density over data streams in a nonstationary environment. In stationary environments, well-established traditional KDE techniques have nice asymptotic properties. Their existing extensions to deal with stream data are mostly based on various heuristic concepts (losing convergence properties). In this paper, we study recursive KDEs, called recursive concept drift tracking KDEs, and prove their weak (in probability) and strong (with probability one) convergence, resulting in perfect tracking properties as the sample size approaches infinity. In three theorems and subsequent examples, we show how to choose the bandwidth and learning rate of a recursive KDE in order to ensure weak and strong convergence. The simulation results illustrate the effectiveness of our algorithm both for density estimation and classification over time-varying stream data. Piotr Duda, Leszek Rutkowski, Maciej Jaworski, Danuta Rutkowska |
IEEE Trans. Cybern. | 1 |
| 2019 | On Handling Missing Values in Data Stream Mining Algorithms Based on the Restricted Boltzmann Machine
Maciej Jaworski, Piotr Duda, Danuta Rutkowska, Leszek Rutkowski |
ICONIP (5) | 2 |
| 2019 | Corrigendum to 'How to adjust an ensemble size in stream data mining?' Information Sciences, vol. 381 (2017), pp. 46-54
Lena Pietruczuk, Leszek Rutkowski, Maciej Jaworski, Piotr Duda |
Inf. Sci. | 4 |
| 2018 | Concept Drift Detection in Streams of Labelled Data Using the Restricted Boltzmann MachineabstractIn this paper, the method of concept drift detection in time-varying data stream mining is considered. The Restricted Boltzmann Machine (RBM) is proposed to be applied as a drift detector. The RBMs which are able to learn joint probability distributions of attribute values and their classes were taken into account. Properly learned they contain a compressed information about the underlying data distribution. The RBM learned on a part of the data stream can be used to determine possible changes in the data stream probability distribution. Two evaluation measures are applied as indicators of possible sudden or gradual changes: the reconstruction error and the free energy. In experiments conducted on synthetic datasets, both measures proved to be well suited for the task of concept drift detection. Maciej Jaworski, Piotr Duda, Leszek Rutkowski |
IJCNN | 2 |
| 2018 | Online GRNN-Based Ensembles for Regression on Evolving Data Streams
Piotr Duda, Maciej Jaworski, Leszek Rutkowski |
ISNN | 1 |
| 2018 | Convergent Time-Varying Regression Models for Data Streams: Tracking Concept Drift by the Recursive Parzen-Based Generalized Regression Neural NetworksabstractOne of the greatest challenges in data mining is related to processing and analysis of massive data streams. Contrary to traditional static data mining problems, data streams require that each element is processed only once, the amount of allocated memory is constant and the models incorporate changes of investigated streams. A vast majority of available methods have been developed for data stream classification and only a few of them attempted to solve regression problems, using various heuristic approaches. In this paper, we develop mathematically justified regression models working in a time-varying environment. More specifically, we study incremental versions of generalized regression neural networks, called IGRNNs, and we prove their tracking properties - weak (in probability) and strong (with probability one) convergence assuming various concept drift scenarios. First, we present the IGRNNs, based on the Parzen kernels, for modeling stationary systems under nonstationary noise. Next, we extend our approach to modeling time-varying systems under nonstationary noise. We present several types of concept drifts to be handled by our approach in such a way that weak and strong convergence holds under certain conditions. Finally, in the series of simulations, we compare our method with commonly used heuristic approaches, based on forgetting mechanism or sliding windows, to deal with concept drift. Finally, we apply our concept in a real life scenario solving the problem of currency exchange rates prediction. Piotr Duda, Maciej Jaworski, Leszek Rutkowski |
Int. J. Neural Syst. | 1 |
| 2018 | Knowledge discovery in data streams with the orthogonal series-based generalized regression neural networks
Piotr Duda, Maciej Jaworski, Leszek Rutkowski |
Inf. Sci. | 1 |
| 2018 | New Splitting Criteria for Decision Trees in Stationary Data StreamsabstractThe most popular tools for stream data mining are based on decision trees. In previous 15 years, all designed methods, headed by the very fast decision tree algorithm, relayed on Hoeffding's inequality and hundreds of researchers followed this scheme. Recently, we have demonstrated that although the Hoeffding decision trees are an effective tool for dealing with stream data, they are a purely heuristic procedure; for example, classical decision trees such as ID3 or CART cannot be adopted to data stream mining using Hoeffding's inequality. Therefore, there is an urgent need to develop new algorithms, which are both mathematically justified and characterized by good performance. In this paper, we address this problem by developing a family of new splitting criteria for classification in stationary data streams and investigating their probabilistic properties. The new criteria, derived using appropriate statistical tools, are based on the misclassification error and the Gini index impurity measures. The general division of splitting criteria into two types is proposed. Attributes chosen based on type- splitting criteria guarantee, with high probability, the highest expected value of split measure. Type- criteria ensure that the chosen attribute is the same, with high probability, as it would be chosen based on the whole infinite data stream. Moreover, in this paper, two hybrid splitting criteria are proposed, which are the combinations of single criteria based on the misclassification error and Gini index. Maciej Jaworski, Piotr Duda, Leszek Rutkowski |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | How to adjust an ensemble size in stream data mining?
Lena Pietruczuk, Leszek Rutkowski, Maciej Jaworski, Piotr Duda |
Inf. Sci. | 4 |
| 2016 | A method for automatic adjustment of ensemble size in stream data miningabstractIn recent years plenty of new algorithms for data stream classification were developed. The occurrence of different concept drift types in data streams turned out to be especially challenging. Much attention was paid to the ensemble methods because of their desired properties. However, the problem of deciding how many components should be stored in the ensemble is still an open issue. Therefore in this article we show a theoretically justified method of determining the proper ensemble size automatically. The performance of the proposed algorithm was experimentally tested and compared with other known methods. Lena Pietruczuk, Leszek Rutkowski, Maciej Jaworski, Piotr Duda |
IJCNN | 4 |
| 2015 | A New Method for Data Stream Mining Based on the Misclassification ErrorabstractIn this paper, a new method for constructing decision trees for stream data is proposed. First a new splitting criterion based on the misclassification error is derived. A theorem is proven showing that the best attribute computed in considered node according to the available data sample is the same, with some high probability, as the attribute derived from the whole infinite data stream. Next this result is combined with the splitting criterion based on the Gini index. It is shown that such combination provides the highest accuracy among all studied algorithms. Leszek Rutkowski, Maciej Jaworski, Lena Pietruczuk, Piotr Duda |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2014 | A novel application of Hoeffding's inequality to decision trees construction for data streamsabstractDecision trees are the commonly applied tools in the task of data stream classification. The most critical point in decision tree construction algorithm is the choice of the splitting attribute. In majority of algorithms existing in literature the splitting criterion is based on statistical bounds derived for split measure functions. In this paper we propose a totally new kind of splitting criterion. We derive statistical bounds for arguments of split measure function instead of deriving it for split measure function itself. This approach allows us to properly use the Hoeffding's inequality to obtain the required bounds. Based on this theoretical results we propose the Decision Trees based on the Fractions Approximation algorithm (DTFA). The algorithm exhibits satisfactory results of classification accuracy in numerical experiments. It is also compared with other existing in literature methods, demonstrating noticeably better performance. Piotr Duda, Maciej Jaworski, Lena Pietruczuk, Leszek Rutkowski |
IJCNN | 1 |
| 2014 | The Parzen kernel approach to learning in non-stationary environmentabstractIn this paper a method for nonparametric regression estimation in non-stationary environment is presented. The Parzen kernels are used to design the recursive general regression neural networks to track changes of non-stationary system under non-stationary noise. The probabilistic properties of the proposed method are investigated. Experimental results are presented and discussed. Lena Pietruczuk, Leszek Rutkowski, Maciej Jaworski, Piotr Duda |
IJCNN | 4 |
| 2014 | The CART decision tree for mining data streams
Leszek Rutkowski, Maciej Jaworski, Lena Pietruczuk, Piotr Duda |
Inf. Sci. | 4 |
| 2014 | Decision Trees for Mining Data Streams Based on the Gaussian ApproximationabstractSince the Hoeffding tree algorithm was proposed in the literature, decision trees became one of the most popular tools for mining data streams. The key point of constructing the decision tree is to determine the best attribute to split the considered node. Several methods to solve this problem were presented so far. However, they are either wrongly mathematically justified (e.g., in the Hoeffding tree algorithm) or time-consuming (e.g., in the McDiarmid tree algorithm). In this paper, we propose a new method which significantly outperforms the McDiarmid tree algorithm and has a solid mathematical basis. Our method ensures, with a high probability set by the user, that the best attribute chosen in the considered node using a finite data sample is the same as it would be in the case of the whole data stream. Leszek Rutkowski, Maciej Jaworski, Lena Pietruczuk, Piotr Duda |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2013 | Decision Trees for Mining Data Streams Based on the McDiarmid's BoundabstractIn mining data streams the most popular tool is the Hoeffding tree algorithm. It uses the Hoeffding's bound to determine the smallest number of examples needed at a node to select a splitting attribute. In the literature the same Hoeffding's bound was used for any evaluation function (heuristic measure), e.g., information gain or Gini index. In this paper, it is shown that the Hoeffding's inequality is not appropriate to solve the underlying problem. We prove two theorems presenting the McDiarmid's bound for both the information gain, used in ID3 algorithm, and for Gini index, used in Classification and Regression Trees (CART) algorithm. The results of the paper guarantee that a decision tree learning system, applied to data streams and based on the McDiarmid's bound, has the property that its output is nearly identical to that of a conventional learner. The results of the paper have a great impact on the state of the art of mining data streams and various developed so far methods and algorithms should be reconsidered. Leszek Rutkowski, Lena Pietruczuk, Piotr Duda, Maciej Jaworski |
IEEE Trans. Knowl. Data Eng. | 3 |