VLDB 2026 Research / reviewers in the wild / expert
Michal Koziarski
dblp:180/4935
· DBLP profile ↗
19ranked-venue papers
12as first author
13since 2021 · last 2025
0000-0001-7707-9640ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 12 first-author · 12 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorTheory of computation · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Action abstractions for amortized samplingabstractAs trajectories sampled by policies used by reinforcement learning (RL) and generative flow networks (GFlowNets) grow longer, credit assignment and exploration become more challenging, and the long planning horizon hinders mode discovery and generalization.
The challenge is particularly pronounced in entropy-seeking RL methods, such as generative flow networks, where the agent must learn to sample from a structured distribution and discover multiple high-reward states, each of which take many steps to reach.
To tackle this challenge, we propose an approach to incorporate the discovery of action abstractions, or high-level actions, into the policy optimization process.
Our approach involves iteratively extracting action subsequences commonly used across many high-reward trajectories and `chunking' them into a single action that is added to the action space.
In empirical evaluation on synthetic and real-world environments, our approach demonstrates improved sample efficiency performance in discovering diverse high-reward objects, especially on harder exploration problems.
We also observe that the abstracted high-order actions are potentially interpretable, capturing the latent structure of the reward landscape of the action space.
This work provides a cognitively motivated approach to action abstraction in RL and is the first demonstration of hierarchical planning in amortized sequential sampling. Oussama Boussif, Léna Néhale Ezzine, Joseph D. Viviano, Michal Koziarski, Moksh Jain, Nikolay Malkin, Emmanuel Bengio, Rim Assouel, Yoshua Bengio |
ICLR | 4 |
| 2025 | Measuring Scientific Capabilities of Language Models with a Systems Biology Dry LababstractDesigning experiments and result interpretations are core scientific competencies, particularly in biology, where researchers perturb complex systems to uncover the underlying systems. Recent efforts to evaluate the scientific capabilities of large language models (LLMs) fail to test these competencies because wet-lab experimentation is prohibitively expensive: in expertise, time and equipment. We introduce SciGym, a first-in-class benchmark that assesses LLMs' iterative experiment design and analysis abilities in open-ended scientific discovery tasks. SciGym overcomes the challenge of wet-lab costs by running a dry lab of biological systems. These models, encoded in Systems Biology Markup Language, are efficient for generating simulated data, making them ideal testbeds for experimentation on realistically complex systems. We evaluated six frontier LLMs on 137 small systems, and released a total of 350 systems at https://huggingface.co/datasets/h4duan/scigym-sbml. Our evaluation shows that while more capable models demonstrated superior performance, all models' performance declined significantly as system complexity increased, suggesting substantial room for improvement in the scientific capabilities of LLM agents. Haonan Duan 0002, Stephen Zhewen Lu, Caitlin F. Harrigan, Nishkrit Desai, Jiarui Lu, Michal Koziarski, Leonardo Cotta, Chris J. Maddison |
NeurIPS | 6 |
| 2025 | Scalable and Cost-Efficient de Novo Template-Based Molecular GenerationabstractTemplate-based molecular generation offers a promising avenue for drug design by ensuring generated compounds are synthetically accessible through predefined reaction templates and building blocks. In this work, we tackle three core challenges in template-based GFlowNets: (1) minimizing synthesis cost, (2) scaling to large building block libraries, and (3) effectively utilizing small fragment sets. We propose **Recursive Cost Guidance**, a backward policy framework that employs auxiliary machine learning models to approximate synthesis cost and viability. This guidance steers generation toward low-cost synthesis pathways, significantly enhancing cost-efficiency, molecular diversity, and quality, especially when paired with an **Exploitation Penalty** that balances the trade-off between exploration and exploitation. To enhance performance in smaller building block libraries, we develop a **Dynamic Library** mechanism that reuses intermediate high-reward states to construct full synthesis trees. Our approach establishes state-of-the-art results in template-based molecular generation. Piotr Gainski, Oussama Boussif, Andrei Rekesh, Dmytro Shevchuk, Ali Parviz, Mike Tyers, Robert A. Batey, Michal Koziarski |
NeurIPS | 8 |
| 2025 | Diverse and feasible retrosynthesis using GFlowNets
Piotr Gainski, Michal Koziarski, Krzysztof Maziarz, Marwin H. S. Segler, Jacek Tabor, Marek Smieja |
Inf. Sci. | 2 |
| 2024 | Towards Foundational Models for Molecular Learning on Large-Scale Multi-Task DatasetsabstractRecently, pre-trained foundation models have enabled significant advancements in multiple fields. In molecular machine learning, however, where datasets are often hand-curated, and hence typically small, the lack of datasets with labeled features, and codebases to manage those datasets, has hindered the development of foundation models. In this work, we present seven novel datasets categorized by size into three distinct categories: ToyMix, LargeMix and UltraLarge. These datasets push the boundaries in both the scale and the diversity of supervised labels for molecular learning. They cover nearly 100 million molecules and over 3000 sparsely defined tasks, totaling more than 13 billion individual labels of both quantum and biological nature. In comparison, our datasets contain 300 times more data points than the widely used OGB-LSC PCQM4Mv2 dataset, and 13 times more than the quantum-only QM1B dataset. In addition, to support the development of foundational models based on our proposed datasets, we present the Graphium graph machine learning library which simplifies the process of building and training molecular machine learning models for multi-task and multi-level molecular datasets. Finally, we present a range of baseline results as a starting point of multi-task and multi-level training on these datasets. Empirically, we observe that performance on low-resource biological datasets show improvement by also training on large amounts of quantum data. This indicates that there may be potential in multi-task and multi-level training of a foundation model and fine-tuning it to resource-constrained downstream tasks. The Graphium library is publicly available on Github and the dataset links are available in Part 1 and Part 2. Dominique Beaini, Shenyang Huang, Joao Alex Cunha, Gabriela Moisescu-Pareja, Oleksandr Dymov, Samuel Maddrell-Mander, Callum McLean, Frederik Wenkel, Luis Müller, Jama Hussein Mohamud, Ali Parviz, Michael Craig, Michal Koziarski, Jiarui Lu, Zhaocheng Zhu, Cristian Gabellini, Kerstin Kläser 0001, Josef Dean, Cas Wognum, Maciej Sypetkowski, Guillaume Rabusseau, Reihaneh Rabbany, Jian Tang 0005, Christopher Morris 0001, Mirco Ravanelli, Guy Wolf, Prudencio Tossou, Hadrien Mary, Therence Bois, Andrew W. Fitzgibbon, Blazej Banaszewski, Chad Martin, Dominic Masters |
ICLR | 14 |
| 2024 | RGFN: Synthesizable Molecular Generation Using GFlowNetsabstractGenerative models hold great promise for small molecule discovery, significantly increasing the size of search space compared to traditional in silico screening libraries. However, most existing machine learning methods for small molecule generation suffer from poor synthesizability of candidate compounds, making experimental validation difficult. In this paper we propose Reaction-GFlowNet (RGFN), an extension of the GFlowNet framework that operates directly in the space of chemical reactions, thereby allowing out-of-the-box synthesizability while maintaining comparable quality of generated candidates. We demonstrate that with the proposed set of reactions and building blocks, it is possible to obtain a search space of molecules orders of magnitude larger than existing screening libraries coupled with low cost of synthesis. We also show that the approach scales to very large fragment libraries, further increasing the number of potential molecules. We demonstrate the effectiveness of the proposed approach across a range of oracle models, including pretrained proxy models and GPU-accelerated docking. Michal Koziarski, Andrei Rekesh, Dmytro Shevchuk, Almer van der Sloot, Piotr Gainski, Yoshua Bengio, Cheng-Hao Liu, Mike Tyers, Robert A. Batey |
NeurIPS | 1 |
| 2024 | Local neighborhood encodings for imbalanced data classificationabstractAbstract This paper aims to propose Local Neighborhood Encodings (LNE)-a hybrid data preprocessing method dedicated to skewed class distribution balancing. The proposed LNE algorithm uses both over- and undersampling methods. The intensity of the methods is chosen separately for each fraction of minority and majority class objects. It is selected depending on the type of neighborhoods of objects of a given class, understood as the number of neighbors from the same class closest to a given object. The process of selecting the over- and undersampling intensities is treated as an optimization problem for which an evolutionary algorithm is used. The quality of the proposed method was evaluated through computer experiments. Compared with SOTA resampling strategies, LNE shows very good results. In addition, an experimental analysis of the algorithms behavior was performed, i.e., the determination of data preprocessing parameters depending on the selected characteristics of the decision problem, as well as the type of classifier used. An ablation study was also performed to evaluate the influence of components on the quality of the obtained classifiers. The evaluation of how the quality of classification is influenced by the evaluation of the objective function in an evolutionary algorithm is presented. In the considered task, the objective function is not de facto deterministic and its value is subject to estimation. Hence, it was important from the point of view of computational efficiency to investigate the possibility of using for quality assessment the so-called proxy classifier, i.e., a classifier of low computational complexity, although the final model was learned using a different model. The proposed data preprocessing method has high quality compared to SOTA, however, it should be noted that it requires significantly more computational effort. Nevertheless, it can be successfully applied to the case as no very restrictive model building time constraints are imposed. Michal Koziarski, Michal Wozniak 0001 |
Mach. Learn. | 1 |
| 2023 | ChiENN: Embracing Molecular Chirality with Graph Neural Networks
Piotr Gainski, Michal Koziarski, Jacek Tabor, Marek Smieja |
ECML/PKDD (3) | 2 |
| 2021 | RB-CCR: Radial-Based Combined Cleaning and Resampling algorithm for imbalanced data classification
Michal Koziarski, Colin Bellinger, Michal Wozniak 0001 |
DSAA | 1 |
| 2021 | CSMOUTE: Combined Synthetic Oversampling and Undersampling Technique for Imbalanced Data ClassificationabstractIn this paper we propose a novel data-level algorithm for handling data imbalance in the classification task, Synthetic Majority Undersampling Technique (SMUTE). SMUTE leverages the concept of interpolation of nearby instances, previously introduced in the oversampling setting in SMOTE. Furthermore, we combine both in the Combined Synthetic Oversampling and Undersampling Technique (CSMOUTE), which integrates SMOTE oversampling with SMUTE undersampling. The results of the conducted experimental study demonstrate the usefulness of both the SMUTE and the CSMOUTE algorithms, especially when combined with more complex classifiers, namely MLP and SVM, and when applied on datasets consisting of a large number of outliers. This leads us to a conclusion that the proposed approach shows promise for further extensions accommodating local data characteristics, a direction discussed in more detail in the paper. Michal Koziarski |
IJCNN | 1 |
| 2021 | Two-Stage Resampling for Convolutional Neural Network Training in the Imbalanced Colorectal Cancer Image ClassificationabstractData imbalance remains one of the open challenges in the contemporary machine learning. It is especially prevalent in case of medical data, such as histopathological images. Traditional data-level approaches for dealing with data imbalance are ill-suited for image data: oversampling methods such as SMOTE and its derivatives lead to creation of unrealistic synthetic observations, whereas undersampling reduces the amount of available data, critical for successful training of convolutional neural networks. To alleviate the problems associated with over- and undersampling we propose a novel two-stage resampling methodology, in which we initially use the oversampling techniques in the image space to leverage a large amount of data for training of a convolutional neural network, and afterwards apply undersampling in the feature space to fine-tune the last layers of the network. Experiments conducted on a colorectal cancer image dataset indicate the usefulness of the proposed approach. Michal Koziarski |
IJCNN | 1 |
| 2021 | RB-CCR: Radial-Based Combined Cleaning and Resampling algorithm for imbalanced data classificationabstractAbstract Real-world classification domains, such as medicine, health and safety, and finance, often exhibit imbalanced class priors and have asynchronous misclassification costs. In such cases, the classification model must achieve a high recall without significantly impacting precision. Resampling the training data is the standard approach to improving classification performance on imbalanced binary data. However, the state-of-the-art methods ignore the local joint distribution of the data or correct it as a post-processing step. This can causes sub-optimal shifts in the training distribution, particularly when the target data distribution is complex. In this paper, we propose Radial-Based Combined Cleaning and Resampling (RB-CCR). RB-CCR utilizes the concept of class potential to refine the energy-based resampling approach of CCR. In particular, RB-CCR exploits the class potential to accurately locate sub-regions of the data-space for synthetic oversampling. The category sub-region for oversampling can be specified as an input parameter to meet domain-specific needs or be automatically selected via cross-validation. Our $$5\times 2$$ 5 × 2 cross-validated results on 57 benchmark binary datasets with 9 classifiers show that RB-CCR achieves a better precision-recall trade-off than CCR and generally out-performs the state-of-the-art resampling methods in terms of AUC and G-mean. Michal Koziarski, Colin Bellinger, Michal Wozniak 0001 |
Mach. Learn. | 1 |
| 2021 | Potential Anchoring for imbalanced data classificationabstractData imbalance remains one of the factors negatively affecting the performance of contemporary machine learning algorithms. One of the most common approaches to reducing the negative impact of data imbalance is preprocessing the original dataset with data-level strategies. In this paper we propose a unified framework for imbalanced data over- and undersampling. The proposed approach utilizes radial basis functions to preserve the original shape of the underlying class distributions during the resampling process. This is done by optimizing the positions of generated synthetic observations with respect to the proposed potential resemblance loss. The final Potential Anchoring algorithm combines over- and undersampling within the proposed framework. The results of the experiments conducted on 60 imbalanced datasets show outperformance of Potential Anchoring over state-of-the-art resampling algorithms, including previously proposed methods that utilize radial basis functions to model class potential. Furthermore, the results of the analysis based on the proposed data complexity index show that Potential Anchoring is particularly well suited for handling naturally complex (i.e. not affected by the presence of noise) datasets. Michal Koziarski |
Pattern Recognit. | 1 |
| 2020 | Combined Cleaning and Resampling algorithm for multi-class imbalanced data with label noise
Michal Koziarski, Michal Wozniak 0001, Bartosz Krawczyk |
Knowl. Based Syst. | 1 |
| 2020 | Radial-Based Undersampling for imbalanced data classificationabstractData imbalance remains one of the most widespread problems affecting contemporary machine learning. The negative effect data imbalance can have on the traditional learning algorithms is most severe in combination with other dataset difficulty factors, such as small disjuncts, presence of outliers and insufficient number of training observations. Aforementioned difficulty factors can also limit the applicability of some of the methods of dealing with data imbalance, in particular the neighborhood-based oversampling algorithms based on SMOTE. Radial-Based Oversampling (RBO) was previously proposed to mitigate some of the limitations of the neighborhood-based methods. In this paper we examine the possibility of utilizing the concept of mutual class potential, used to guide the oversampling process in RBO, in the undersampling procedure. Conducted computational complexity analysis indicates a significantly reduced time complexity of the proposed Radial-Based Undersampling algorithm, and the results of the performed experimental study indicate its usefulness, especially on difficult datasets. Michal Koziarski |
Pattern Recognit. | 1 |
| 2020 | Radial-Based Oversampling for Multiclass Imbalanced Data ClassificationabstractLearning from imbalanced data is among the most popular topics in the contemporary machine learning. However, the vast majority of attention in this field is given to binary problems, while their much more difficult multiclass counterparts are relatively unexplored. Handling data sets with multiple skewed classes poses various challenges and calls for a better understanding of the relationship among classes. In this paper, we propose multiclass radial-based oversampling (MC-RBO), a novel data-sampling algorithm dedicated to multiclass problems. The main novelty of our method lies in using potential functions for generating artificial instances. We take into account information coming from all of the classes, contrary to existing multiclass oversampling approaches that use only minority class characteristics. The process of artificial instance generation is guided by exploring areas where the value of the mutual class distribution is very small. This way, we ensure a smart oversampling procedure that can cope with difficult data distributions and alleviate the shortcomings of existing methods. The usefulness of the MC-RBO algorithm is evaluated on the basis of extensive experimental study and backed-up with a thorough statistical analysis. Obtained results show that by taking into account information coming from all of the classes and conducting a smart oversampling, we can significantly improve the process of learning from multiclass imbalanced data. Bartosz Krawczyk, Michal Koziarski, Michal Wozniak 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | Radial-Based oversampling for noisy imbalanced data classification
Michal Koziarski, Bartosz Krawczyk, Michal Wozniak 0001 |
Neurocomputing | 1 |
| 2017 | The deterministic subspace method for constructing classifier ensemblesabstractEnsemble classification remains one of the most popular techniques in contemporary machine learning, being characterized by both high efficiency and stability. An ideal ensemble comprises mutually complementary individual classifiers which are characterized by the high diversity and accuracy. This may be achieved, e.g., by training individual classification models on feature subspaces. Random Subspace is the most well-known method based on this principle. Its main limitation lies in stochastic nature, as it cannot be considered as a stable and a suitable classifier for real-life applications. In this paper, we propose an alternative approach, Deterministic Subspace method, capable of creating subspaces in guided and repetitive manner. Thus, our method will always converge to the same final ensemble for a given dataset. We describe general algorithm and three dedicated measures used in the feature selection process. Finally, we present the results of the experimental study, which prove the usefulness of the proposed method. Michal Koziarski, Bartosz Krawczyk, Michal Wozniak 0001 |
Pattern Anal. Appl. | 1 |
| 2016 | Forming Classifier Ensembles with Deterministic Feature SubspacesabstractEnsemble learning is being considered as one of the most well-established and efficient techniques in the contemporary machine learning.The key to the satisfactory performance of such combined models lies in the supplied base learners and selected combination strategy.In this paper we will focus on the former issue.Having classifiers that are of high individual quality and complementary to each other is a desirable property.Among several ways to ensure diversity feature space division deserves attention.The most popular method employed here is Random Subspace approach.However, due to its random nature one cannot consider this approach as stable one or suitable for reallife applications.Therefore, we propose a new approach called Deterministic Subspace that constructs feature subspaces in a guided and repetitive manner.We present a general framework and three dedicated measures that can be used for selecting diverse and uncorrelated features for each base learner.This way we will always obtain identical sets of features, leading to creation of stable ensembles.Experimental study backed-up with statistical analysis prove the usefulness of our method in comparison to popular randomized solution. Michal Koziarski, Bartosz Krawczyk, Michal Wozniak 0001 |
FedCSIS | 1 |