VLDB 2026 Research / reviewers in the wild / expert
Jan Mielniczuk
dblp:69/10843
· DBLP profile ↗
14ranked-venue papers
6as first author
10since 2021 · last 2025
0000-0003-2621-2303ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Theory of computation · 2 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Generalized Approach to Label Shift: The Conditional Probability Shift ModelabstractIn many practical applications of machine learning, a discrepancy often arises between a source distribution from which labeled training examples are drawn and a target distribution for which only unlabeled data is observed. Traditionally, two main scenarios have been considered to address this issue: covariate shift (CS), where only the marginal distribution of features changes, and label shift (LS), which involves a change in the class variable’s prior distribution. However, these frameworks do not encompass all forms of distributional shift. This paper introduces a new setting, Conditional Probability Shift (CPS), which captures the case when the conditional distribution of the class variable given some specific features changes while the distribution of remaining features given the specific features and the class is preserved. For this scenario we present the Conditional Probability Shift Model (CPSM) based on modeling the class variable’s conditional probabilities using multinomial regression. Since the class variable is not observed for the target data, the parameters of the multinomial model for its distribution are estimated using the Expectation-Maximization algorithm. The proposed method is generic and can be combined with any probabilistic classifier. The effectiveness of CPSM is demonstrated through experiments on synthetic datasets and a case study using the MIMIC medical database, revealing its superior balanced classification accuracy on the target data compared to existing methods, particularly in situations of conditional distribution shift and no prior distribution shift, which are not detected by LS-based methods. Pawel Teisseyre, Jan Mielniczuk |
ECAI | 2 |
| 2025 | Single-sample Versus Case-control Sampling Scheme for Positive Unlabeled Data: the Story of Two ScenariosabstractIn the paper we argue that performance of the classifiers based on Empirical Risk Minimization (ERM) for positive unlabeled data, which are designed for case-control sampling scheme may significantly deteriorate when applied to a single-sample scenario. We reveal why their behavior depends, in all but very specific cases, on the scenario. Also, we introduce a single-sample case analogue of the popular non-negative risk classifier designed for case-control data and compare its performance with the original proposal. We show that the significant differences occur between them, especially when half or more positive of observations are labeled. The opposite case when ERM minimizer designed for the case-control case is applied for single-sample data is also considered and similar conclusions are drawn. Taking into account difference of scenarios requires a sole, but crucial, change in the definition of the Empirical Risk. Jan Mielniczuk, Adam Wawrzenczyk |
Fundam. Informaticae | 1 |
| 2024 | Augmented Prediction of a True Class for Positive Unlabeled Data Under Selection BiasabstractWe introduce a new observational setting for Positive Unlabeled (PU) data where the observations at prediction time are also labeled. This occurs commonly in practice – we argue that the additional information is important for prediction, and call this task “augmented PU prediction”. We allow for labeling to be feature dependent. In such scenario, Bayes classifier and its risk is established and compared with a risk of a classifier which for unlabeled data is based only on predictors. We introduce several variants of the empirical Bayes rule in such scenario and investigate their performance. We put a special focus on dangers (and ease) of applying classical classification rule in the augmented PU scenario – due to no preexisting studies, an unaware researcher is prone to skewing the obtained predictions. We conclude that the variant based on recently proposed variational autoencoder designed for PU scenario works on par or better than other considered variants and yields advantage over feature-only based methods in terms of accuracy for unlabeled samples. Jan Mielniczuk, Adam Wawrzenczyk |
ECAI | 1 |
| 2024 | Verifying the Selected Completely at Random Assumption in Positive-Unlabeled LearningabstractThe goal of positive-unlabeled (PU) learning is to train a binary classifier on the basis of training data containing positive and unlabeled instances, where unlabeled observations can belong either to the positive class or to the negative class. Modeling PU data requires certain assumptions on the labeling mechanism that describes which positive observations are assigned a label. The simplest assumption, considered in early works, is SCAR (Selected Completely at Random Assumption), according to which the propensity score function, defined as the probability of assigning a label to a positive observation, is constant. Alternatively, a much more realistic assumption is SAR (Selected at Random), which states that the propensity function solely depends on the observed feature vector. SCAR-based algorithms are much simpler and computationally much faster compared to SAR-based algorithms, which usually require challenging estimation of the propensity score. In this work, we propose a relatively simple and computationally fast test that can be used to determine whether the observed data meet the SCAR assumption. Our test is based on generating artificial labels conforming to the SCAR scenario, which in turn allows to mimic the distribution of the test statistic under the null hypothesis of SCAR. We justify our method theoretically. In experiments, we demonstrate that the test successfully detects various deviations from SCAR scenario and at the same time it is possible to effectively control the type I error. The proposed test can be recommended as a pre-processing step to decide which final PU algorithm to choose in cases when nature of labeling mechanism is not known. Pawel Teisseyre, Konrad Furmanczyk, Jan Mielniczuk |
ECAI | 3 |
| 2024 | Joint empirical risk minimization for instance-dependent positive-unlabeled data
Wojciech Rejchel, Pawel Teisseyre, Jan Mielniczuk |
Knowl. Based Syst. | 3 |
| 2023 | Double Logistic Regression Approach to Biased Positive-Unlabeled DataabstractPositive and unlabelled learning is an important non-standard inference problem which arises naturally in many applications. The significant limitation of almost all existing methods addressing it lies in assuming that the propensity score function is constant and does not depend on features (Selected Completely at Random assumption), which is unrealistic in many practical situations. Avoiding this assumption, we consider parametric approach to the problem of joint estimation of posterior probability and propensity score functions. We show that if both these functions are logistic with different parameters (double logistic model) then the corresponding parameters are identifiable. Motivated by this, we propose two approaches to their estimation: a joint maximum likelihood method and the second approach based on an alternating maximization of two Fisher consistent approximations. Our experimental results show that the proposed methods perform on par or better than the existing methods based on Expectation-Maximisation scheme. Konrad Furmanczyk, Jan Mielniczuk, Wojciech Rejchel, Pawel Teisseyre |
ECAI | 2 |
| 2023 | One-Class Classification Approach to Variational Learning from Biased Positive Unlabeled DataabstractWe discuss Empirical Risk Minimization approach in conjunction with one-class classification method to learn classifiers for biased Positive Unlabeled (PU) data. For such data, probability that an observation from a positive class is labeled may depend on its features. The proposed method extends Variational Autoencoder for PU data (VAE-PU) introduced in [16] by proposing another estimator of a theoretical risk of a classifier to be minimized, which has important advantages over the previous proposal. This is based on one-class classification approach using generated pseudo-observations, which turns out to be an effective method of detecting positive observations among unlabeled ones. The proposed method leads to more precise estimation of the theoretical risk than the previous proposal. Experiments performed on real data sets show that the proposed VAE-PU+OCC algorithm works very promisingly in comparison to its competitors such as the original VAE-PU, SAR-EM and LBE methods in terms of accuracy and F1 score. The advantage is especially strongly pronounced for small labeling frequencies. Jan Mielniczuk, Adam Wawrzenczyk |
ECAI | 1 |
| 2023 | Enhancing naive classifier for positive unlabeled data based on logistic regression approachabstractIt is argued that for analysis of Positive Unlabeled (PU) data under Selected Completely At Random (SCAR) assumption it is fruitful to view the problem as fitting of misspecified model to the data.Namely, it is shown that the results on misspecified fit imply that in the case when posterior probability of the response is modelled by logistic regression, fitting the logistic regression to the observable PU data which does not follow this model, still yields the vector of estimated parameters approximately colinear with the true vector of parameters.This observation together with choosing the intercept of the classifier based on optimisation of analogue of F1 measure yields a classifier which performs on par or better than its competitors on several real data sets considered. Mateusz Platek, Jan Mielniczuk |
FedCSIS | 2 |
| 2021 | Multiple Testing of Conditional Independence Hypotheses Using Information-Theoretic Approach
Malgorzata Lazecka, Jan Mielniczuk |
MDAI | 2 |
| 2021 | How to Gain on Power: Novel Conditional Independence Tests Based on Short Expansion of Conditional Mutual InformationabstractConditional independence tests play a crucial role in many machine learning procedures such as feature selection, causal discovery, and structure learning of dependence networks. They are used in most of the existing algorithms for Markov Blanket discovery such as Grow-Shrink or Incremental Association Markov Blanket. One of the most frequently used tests for categorical variables is based on the conditional mutual information ($CMI$) and its asymptotic distribution. However, it is known that the power of such test dramatically decreases when the size of the conditioning set grows, i.e. the test fails to detect true significant variables, when the set of already selected variables is large. To overcome this drawback for discrete data, we propose to replace the conditional mutual information by Short Expansion of Conditional Mutual Information (called $SECMI$), obtained by truncating the Möbius representation of $CMI$. We prove that the distribution of $SECMI$ converges to either a normal distribution or to a distribution of some quadratic form in normal random variables. This property is crucial for the construction of a novel test of conditional independence which uses one of these distributions, chosen in a data dependent way, as a reference under the null hypothesis. The proposed methods have significantly larger power for discrete data than the standard asymptotic tests of conditional independence based on $CMI$ while retaining control of the probability of type I error. Mariusz Kubkowski, Jan Mielniczuk, Pawel Teisseyre |
J. Mach. Learn. Res. | 2 |
| 2019 | Stopping rules for mutual information-based feature selection
Jan Mielniczuk, Pawel Teisseyre |
Neurocomputing | 1 |
| 2015 | Combined l1 and greedy l0 penalized least squares for linear model selection
Piotr Pokarowski, Jan Mielniczuk |
J. Mach. Learn. Res. | 2 |
| 2007 | Decorrelation of Wavelet Coefficients for Long-Range Dependent ProcessesabstractWe consider a discrete-time stationary long-range dependent process$(X_k)_{k\in Z}$such that its spectral density equals${\varphi}(\vert {\lambda}\vert)^{-2d}$, where${\varphi}$is a smooth function such that${\varphi}(0)={\varphi}^{\prime\prime}(0)=0$and${\varphi}({\lambda})\geq c{\lambda}$for${\lambda}\in [0,\pi]$. Then for any wavelet$\psi$with$N$vanishing moments, the lag$k$within-level covariance of wavelet coefficients decays as${\cal O}(k^{2d-2N-1})$when$k\to\infty$. The result applies to fractionally integrated autoregressive moving average (ARMA) processes as well as to fractional Gaussian noise. Jan Mielniczuk, Piotr Wojdyllo |
IEEE Trans. Inf. Theory | 1 |
| 1993 | Consistency of multilayer perceptron regression estimators
Jan Mielniczuk, Joanna Tyrcha |
Neural Networks | 1 |