Jakub M. Tomczak

dblp:80/8238 · also Jakub Mikolaj Tomczak · DBLP profile ↗
← Back
42ranked-venue papers
8as first author
15since 2021 · last 2026
0000-0001-8634-694XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 37 · 6 first-author · 14 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 A Parallel U-Net Concept for Image Recontextualization
Aleksander Skorupa, Jakub Balicki, Konrad Karanowski, Tomasz Drózdz, Tomasz Halas, Hong Lyu, Oriol Caudevilla, Jakub M. Tomczak, Maciej Zieba
ACIIDS (1)8
2025 Knowledge graph-extended retrieval augmented generation for question answering
abstract
Abstract Large Language Models (LLMs) and Knowledge Graphs (KGs) offer a promising approach to robust and explainable Question Answering (QA). While LLMs excel at natural language understanding, they suffer from knowledge gaps and hallucinations. KGs provide structured knowledge but lack natural language interaction. Ideally, an AI system should be both robust to missing facts as well as easy to communicate with. This paper proposes such a system that integrates LLMs and KGs without requiring training, ensuring adaptability across different KGs with minimal human effort. The resulting approach can be classified as a specific form of a Retrieval Augmented Generation (RAG) with a KG, thus, it is dubbed Knowledge Graph-extended Retrieval Augmented Generation (KG-RAG). It includes a question decomposition module to enhance multi-hop information retrieval and answer explainability. Using In-Context Learning (ICL) and Chain-of-Thought (CoT) prompting, it generates explicit reasoning chains processed separately to improve truthfulness. Experiments on the MetaQA benchmark show increased accuracy for multi-hop questions, though with a slight trade-off in single-hop performance compared to LLM with KG baselines. These findings demonstrate KG-RAG’s potential to improve transparency in QA by bridging unstructured language understanding with structured knowledge retrieval.
Jasper Linders, Jakub M. Tomczak
Appl. Intell.2
2024 Mixed Models with Multiple Instance Learning
Jan P. Engelmann, Alessandro Palma, Jakub M. Tomczak, Fabian J. Theis, Francesco Paolo Casale
AISTATS3
2023 Modelling Long Range Dependencies in $N$D: From Task-Specific to a General Purpose CNN
David M. Knigge, David W. Romero, Albert Gu, Efstratios Gavves, Erik J. Bekkers, Jakub M. Tomczak, Mark Hoogendoorn, Jan-Jakob Sonke
ICLR6
2023 A-NeSI: A Scalable Approximate Method for Probabilistic Neurosymbolic Inference
abstract
We study the problem of combining neural networks with symbolic reasoning. Recently introduced frameworks for Probabilistic Neurosymbolic Learning (PNL), such as DeepProbLog, perform exponential-time exact inference, limiting the scalability of PNL solutions. We introduce Approximate Neurosymbolic Inference (A-NeSI): a new framework for PNL that uses neural networks for scalable approximate inference. A-NeSI 1) performs approximate inference in polynomial time without changing the semantics of probabilistic logics; 2) is trained using data generated by the background knowledge; 3) can generate symbolic explanations of predictions; and 4) can guarantee the satisfaction of logical constraints at test time, which is vital in safety-critical applications. Our experiments show that A-NeSI is the first end-to-end method to solve three neurosymbolic tasks with exponential combinatorial scaling. Finally, our experiments show that A-NeSI achieves explainability and safety without a penalty in performance.
Emile van Krieken, Thiviyan Thanapalasingam, Jakub M. Tomczak, Frank van Harmelen, Annette ten Teije
NeurIPS3
2023 Learning Data Representations with Joint Diffusion Models
Kamil Deja, Tomasz Trzcinski, Jakub M. Tomczak
ECML/PKDD (2)3
2022 FlexConv: Continuous Kernel Convolutions With Differentiable Kernel Sizes
David W. Romero, Robert-Jan Bruintjes, Jakub M. Tomczak, Erik J. Bekkers, Mark Hoogendoorn, Jan C. van Gemert
ICLR3
2022 CKConv: Continuous Kernel Convolution For Sequential Data
David W. Romero, Anna Kuzina, Erik J. Bekkers, Jakub M. Tomczak, Mark Hoogendoorn
ICLR4
2022 On Analyzing Generative and Denoising Capabilities of Diffusion-based Deep Generative Models
abstract
Diffusion-based Deep Generative Models (DDGMs) offer state-of-the-art performance in generative modeling. Their main strength comes from their unique setup in which a model (the backward diffusion process) is trained to reverse the forward diffusion process, which gradually adds noise to the input signal. Although DDGMs are well studied, it is still unclear how the small amount of noise is transformed during the backward diffusion process. Here, we focus on analyzing this problem to gain more insight into the behavior of DDGMs and their denoising and generative capabilities. We observe a fluid transition point that changes the functionality of the backward diffusion process from generating a (corrupted) image from noise to denoising the corrupted image to the final sample. Based on this observation, we postulate to divide a DDGM into two parts: a denoiser and a generator. The denoiser could be parameterized by a denoising auto-encoder, while the generator is a diffusion-based model with its own set of parameters. We experimentally validate our proposition, showing its pros and cons.
Kamil Deja, Anna Kuzina, Tomasz Trzcinski, Jakub M. Tomczak
NeurIPS4
2022 Alleviating Adversarial Attacks on Variational Autoencoders with MCMC
abstract
Variational autoencoders (VAEs) are latent variable models that can generate complex objects and provide meaningful latent representations. Moreover, they could be further used in downstream tasks such as classification. As previous work has shown, one can easily fool VAEs to produce unexpected latent representations and reconstructions for a visually slightly modified input. Here, we examine several objective functions for adversarial attacks construction proposed previously and present a solution to alleviate the effect of these attacks. Our method utilizes the Markov Chain Monte Carlo (MCMC) technique in the inference step that we motivate with a theoretical analysis. Thus, we do not incorporate any extra costs during training and the performance on non-attacked inputs is not decreased. We validate our approach on a variety of datasets (MNIST, Fashion MNIST, Color MNIST, CelebA) and VAE configurations ($\beta$-VAE, NVAE, $\beta$-TCVAE), and show that our approach consistently improves the model robustness to adversarial attacks.
Anna Kuzina, Max Welling, Jakub M. Tomczak
NeurIPS3
2021 Selecting Data Augmentation for Simulating Interventions
abstract
Machine learning models trained with purely observational data and the principle of empirical risk minimization (Vapnik 1992) can fail to generalize to unseen domains. In this paper, we focus on the case where the problem arises through spurious correlation between the observed domains and the actual task labels. We find that many domain generalization methods do not explicitly take this spurious correlation into account. Instead, especially in more application-oriented research areas like medical imaging or robotics, data augmentation techniques that are based on heuristics are used to learn domain invariant features. To bridge the gap between theory and practice, we develop a causal perspective on the problem of domain generalization. We argue that causal concepts can be used to explain the success of data augmentation by describing how they can weaken the spurious correlation between the observed domains and the task labels. We demonstrate that data augmentation can serve as a tool for simulating interventional data. We use these theoretical insights to derive a simple algorithm that is able to select data augmentation techniques that will lead to better domain generalization.
Maximilian Ilse, Jakub M. Tomczak, Patrick Forré
ICML2
2021 Storchastic: A Framework for General Stochastic Automatic Differentiation
abstract
Modelers use automatic differentiation (AD) of computation graphs to implement complex Deep Learning models without defining gradient computations. Stochastic AD extends AD to stochastic computation graphs with sampling steps, which arise when modelers handle the intractable expectations common in Reinforcement Learning and Variational Inference. However, current methods for stochastic AD are limited: They are either only applicable to continuous random variables and differentiable functions, or can only use simple but high variance score-function estimators. To overcome these limitations, we introduce Storchastic, a new framework for AD of stochastic computation graphs. Storchastic allows the modeler to choose from a wide variety of gradient estimation methods at each sampling step, to optimally reduce the variance of the gradient estimates. Furthermore, Storchastic is provably unbiased for estimation of any-order gradients, and generalizes variance reduction techniques to higher-order gradient estimates. Finally, we implement Storchastic as a PyTorch library at github.com/HEmile/storchastic.
Emile van Krieken, Jakub M. Tomczak, Annette ten Teije
NeurIPS2
2021 Invertible DenseNets with Concatenated LipSwish
abstract
We introduce Invertible Dense Networks (i-DenseNets), a more parameter efficient extension of Residual Flows. The method relies on an analysis of the Lipschitz continuity of the concatenation in DenseNets, where we enforce invertibility of the network by satisfying the Lipschitz constant. Furthermore, we propose a learnable weighted concatenation, which not only improves the model performance but also indicates the importance of the concatenated weighted representation. Additionally, we introduce the Concatenated LipSwish as activation function, for which we show how to enforce the Lipschitz condition and which boosts performance. The new architecture, i-DenseNet, out-performs Residual Flow and other flow-based models on density estimation evaluated in bits per dimension, where we utilize an equal parameter budget. Moreover, we show that the proposed model out-performs Residual Flows when trained as a hybrid model where the model is both a generative and a discriminative model.
Yura Perugachi-Diaz, Jakub M. Tomczak, Sandjai Bhulai
NeurIPS2
2021 M2R: a Python add-on to cobrapy for modifying human genome-scale metabolic reconstruction using the gut microbiota models
abstract
MOTIVATION: The gut microbiota is the human body's largest population of microorganisms that interact with human intestinal cells. They use ingested nutrients for fundamental biological processes and have important impacts on human physiology, immunity and metabolome in the gastrointestinal tract. RESULTS: Here, we present M2R, a Python add-on to cobrapy that allows incorporating information about the gut microbiota metabolism models to human genome-scale metabolic models (GEMs) like RECON3D. The idea behind the software is to modify the lower bounds of the exchange reactions in the model using aggregated in- and out-fluxes from selected microbes. M2R enables users to quickly and easily modify the pool of the metabolites that enter and leave the GEM, which is particularly important for those looking into an analysis of the metabolic interaction between the gut microbiota and human cells and its dysregulation. AVAILABILITY AND IMPLEMENTATION: M2R is freely available under an MIT License at https://github.com/e-weglarz-tomczak/m2r. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ewelina Weglarz-Tomczak, Jakub M. Tomczak, Stanley Brul
Bioinform.2
2021 Learning locomotion skills in evolvable robots
abstract
The challenge of robotic reproduction – making of new robots by recombining two existing ones – has been recently cracked and physically evolving robot systems have come within reach. Here we address the next big hurdle: producing an adequate brain for a newborn robot. In particular, we address the task of targeted locomotion which is arguably a fundamental skill in any practical implementation. We introduce a controller architecture and a generic learning method to allow a modular robot with an arbitrary shape to learn to walk towards a target and follow this target if it moves. Our approach is validated on three robots, a spider, a gecko, and their offspring, in three real-world scenarios.
Gongjin Lan, Maarten van Hooft, Matteo De Carlo, Jakub M. Tomczak, A. E. Eiben
Neurocomputing4
2020 Evolutionary Algorithm with Non-parametric Surrogate Model for Tensor Program optimization
abstract
The efficiency of tensor operators is key to implement fast deep learning models. However, identifying the fastest implementation of a tensor operator for a target hardware is challenging. A wide range of different configurations have to be considered, and the evaluation of a configuration is time consuming as it requires compilation and execution of the operator. A common approach to address these issues is to boost traditional optimization algorithms with a surrogate modet, i.e., a machine learning model that approximates the objective function and is cheap to query compared to the target hardware. However, as the surrogate model grows in complexity, so does the time needed to train and maintain it. In this work, we propose to use an evolutionary optimizer and augment it with a non-parametric surrogate model (a weighted k-Nearest-Neighbor regression). We evaluate our approach on the convolution layers of a ResNetl8, and show a convergence speedup of up to 1.4×; when compared to baseline operator tuners.
Ioannis Gatopoulos, Romain Lepert, Auke J. Wiggers, Giovanni Mariani, Jakub M. Tomczak
CEC5
2020 Conditional Channel Gated Networks for Task-Aware Continual Learning
abstract
Convolutional Neural Networks experience catastrophic forgetting when optimized on a sequence of learning problems: as they meet the objective of the current training examples, their performance on previous tasks drops drastically. In this work, we introduce a novel framework to tackle this problem with conditional computation. We equip each convolutional layer with task-specific gating modules, selecting which filters to apply on the given input. This way, we achieve two appealing properties. Firstly, the execution patterns of the gates allow to identify and protect important filters, ensuring no loss in the performance of the model for previously learned tasks. Secondly, by using a sparsity objective, we can promote the selection of a limited set of kernels, allowing to retain sufficient model capacity to digest new tasks. Existing solutions require, at test time, awareness of the task to which each example belongs to. This knowledge, however, may not be available in many practical scenarios. Therefore, we additionally introduce a task classifier that predicts the task label of each example, to deal with settings in which a task oracle is not available. We validate our proposal on four continual learning datasets. Results show that our model consistently outperforms existing methods both in the presence and the absence of a task oracle. Notably, on Split SVHN and Imagenet-50 datasets, our model yields up to 23.98% and 17.42% improvement in accuracy w.r.t. competing methods.
Davide Abati, Jakub M. Tomczak, Tijmen Blankevoort, Simone Calderara, Rita Cucchiara, Babak Ehteshami Bejnordi
CVPR2
2020 Attentive Group Equivariant Convolutional Networks
abstract
Although group convolutional networks are able to learn powerful representations based on symmetry patterns, they lack explicit means to learn meaningful relationships among them (e.g., relative positions and poses). In this paper, we present attentive group equivariant convolutions, a generalization of the group convolution, in which attention is applied during the course of convolution to accentuate meaningful symmetry combinations and suppress non-plausible, misleading ones. We indicate that prior work on visual attention can be described as special cases of our proposed framework and show empirically that our attentive group equivariant convolutional networks consistently outperform conventional group convolutional networks on benchmark image datasets. Simultaneously, we provide interpretability to the learned concepts through the visualization of equivariant attention maps.
David W. Romero, Erik J. Bekkers, Jakub M. Tomczak, Mark Hoogendoorn
ICML3
2020 The Convolution Exponential and Generalized Sylvester Flows
abstract
This paper introduces a new method to build linear flows, by taking the exponential of a linear transformation. This linear transformation does not need to be invertible itself, and the exponential has the following desirable properties: it is guaranteed to be invertible, its inverse is straightforward to compute and the log Jacobian determinant is equal to the trace of the linear transformation. An important insight is that the exponential can be computed implicitly, which allows the use of convolutional layers. Using this insight, we develop new invertible transformations named convolution exponentials and graph convolution exponentials, which retain the equivariance of their underlying transformations. In addition, we generalize Sylvester Flows and propose Convolutional Sylvester Flows which are based on the generalization and the convolution exponential as basis change. Empirically, we show that the convolution exponential outperforms other linear transformations in generative flows on CIFAR10 and the graph convolution exponential improves the performance of graph normalizing flows. In addition, we show that Convolutional Sylvester Flows improve performance over residual flows as a generative flow model measured in log-likelihood.
Emiel Hoogeboom, Victor Garcia Satorras, Jakub M. Tomczak, Max Welling
NeurIPS3
2019 Video Compression With Rate-Distortion Autoencoders
abstract
In this paper we present a deep generative model for lossy video compression. We employ a model that consists of a 3D autoencoder with a discrete latent space and an autoregressive prior used for entropy coding. Both autoencoder and prior are trained jointly to minimize a ratedistortion loss, which is closely related to the ELBO used in variational autoencoders. Despite its simplicity, we find that our method outperforms the state-of-the-art learned video compression networks based on motion compensation or interpolation. We systematically evaluate various design choices, such as the use offrame-based or spatio-temporal autoencoders, and the type of autoregressive prior. In addition, we present three extensions of the basic method that demonstrate the benefits over classical approaches to compression. First, we introduce semantic compression, where the model is trained to allocate more bits to objects of interest. Second, we study adaptive compression, where the model is adapted to a domain with limited variability, e.g. videos taken from an autonomous car, to achieve superior compression on that domain. Finally, we introduce multimodal compression, where we demonstrate the effectiveness of our model in joint compression of multiple modalities captured by non-standard imaging sensors, such as quad cameras. We believe that this opens up novel video compression applications, which have not been feasible with classical codecs.
AmirHossein Habibian, Ties van Rozendaal, Jakub M. Tomczak, Taco Cohen
ICCV3
2019 Combinatorial Bayesian Optimization using the Graph Cartesian Product
abstract
This paper focuses on Bayesian Optimization (BO) for objectives on combinatorial search spaces, including ordinal and categorical variables. Despite the abundance of potential applications of Combinatorial BO, including chipset configuration search and neural architecture search, only a handful of methods have been pro- posed. We introduce COMBO, a new Gaussian Process (GP) BO. COMBO quantifies “smoothness” of functions on combinatorial search spaces by utilizing a combinatorial graph. The vertex set of the combinatorial graph consists of all possible joint assignments of the variables, while edges are constructed using the graph Cartesian product of the sub-graphs that represent the individual variables. On this combinatorial graph, we propose an ARD diffusion kernel with which the GP is able to model high-order interactions between variables leading to better performance. Moreover, using the Horseshoe prior for the scale parameter in the ARD diffusion kernel results in an effective variable selection procedure, making COMBO suitable for high dimensional problems. Computationally, in COMBO the graph Cartesian product allows the Graph Fourier Transform calculation to scale linearly instead of exponentially.We validate COMBO in a wide array of real- istic benchmarks, including weighted maximum satisfiability problems and neural architecture search. COMBO outperforms consistently the latest state-of-the-art while maintaining computational and statistical efficiency
ChangYong Oh, Jakub M. Tomczak, Efstratios Gavves, Max Welling
NeurIPS2
2019 Low-Dimensional Perturb-and-MAP Approach for Learning Restricted Boltzmann Machines
abstract
This paper introduces a new approach to maximum likelihood learning of the parameters of a restricted Boltzmann machine (RBM). The proposed method is based on the Perturb-and-MAP (PM) paradigm that enables sampling from the Gibbs distribution. PM is a two step process: (i) perturb the model using Gumbel perturbations, then (ii) find the maximum a posteriori (MAP) assignment of the perturbed model. We show that under certain conditions the resulting MAP configuration of the perturbed model is an unbiased sample from the original distribution. However, this approach requires an exponential number of perturbations, which is computationally intractable. Here, we apply an approximate approach based on the first order (low-dimensional) PM to calculate the gradient of the log-likelihood in binary RBM. Our approach relies on optimizing the energy function with respect to observable and hidden variables using a greedy procedure. First, for each variable we determine whether flipping this value will decrease the energy, and then we utilize the new local maximum to approximate the gradient. Moreover, we show that in some cases our approach works better than the standard coordinate-descent procedure for finding the MAP assignment and compare it with the Contrastive Divergence algorithm. We investigate the quality of our approach empirically, first on toy problems, then on various image datasets and a text dataset.
Jakub M. Tomczak, Szymon Zareba, Siamak Ravanbakhsh, Russell Greiner
Neural Process. Lett.1
2018 VAE with a VampPrior
abstract
Many different methods to train deep generative models have been introduced in the past. In this paper, we propose to extend the variational auto-encoder (VAE) framework with a new type of prior which we call "Variational Mixture of Posteriors" prior, or VampPrior for short. The VampPrior consists of a mixture distribution (e.g., a mixture of Gaussians) with components given by variational posteriors conditioned on learnable pseudo-inputs. We further extend this prior to a two layer hierarchical model and show that this architecture with a coupled prior and posterior, learns significantly better models. The model also avoids the usual local optima issues related to useless latent dimensions that plague VAEs. We provide empirical studies on six datasets, namely, static and binary MNIST, OMNIGLOT, Caltech 101 Silhouettes, Frey Faces and Histopathology patches, and show that applying the hierarchical VampPrior delivers state-of-the-art results on all datasets in the unsupervised permutation invariant setting and the best results or comparable to SOTA methods for the approach with convolutional networks.
Jakub M. Tomczak, Max Welling
AISTATS1
2018 Attention-based Deep Multiple Instance Learning
abstract
Multiple instance learning (MIL) is a variation of supervised learning where a single class label is assigned to a bag of instances. In this paper, we state the MIL problem as learning the Bernoulli distribution of the bag label where the bag label probability is fully parameterized by neural networks. Furthermore, we propose a neural network-based permutation-invariant aggregation operator that corresponds to the attention mechanism. Notably, an application of the proposed attention-based operator provides insight into the contribution of each instance to the bag label. We show empirically that our approach achieves comparable performance to the best MIL methods on benchmark MIL datasets and it outperforms other methods on a MNIST-based MIL dataset and two real-life histopathology datasets without sacrificing interpretability.
Maximilian Ilse, Jakub M. Tomczak, Max Welling
ICML2
2018 Sylvester Normalizing Flows for Variational Inference
Rianne van den Berg, Leonard Hasenclever, Jakub M. Tomczak, Max Welling
UAI3
2018 Hyperspherical Variational Auto-Encoders
Tim R. Davidson, Luca Falorsi, Nicola De Cao, Thomas Kipf, Jakub M. Tomczak
UAI5
2017 Learning Invariant Features Using Subspace Restricted Boltzmann Machine
abstract
The subspace restricted Boltzmann machine (subspaceRBM) is a third-order Boltzmann machine where multiplicative interactions are between one visible and two hidden units. There are two kinds of hidden units, namely, gate units and subspace units . The subspace units reflect variations of a pattern in data and the gate unit is responsible for activating the subspace units. Additionally, the gate unit can be seen as a pooling feature. We evaluate the behavior of subspaceRBM through experiments with MNIST digit recognition task and Caltech 101 Silhouettes image corpora, measuring cross-entropy reconstruction error and classification error.
Jakub M. Tomczak, Adam Gonczarek
Neural Process. Lett.1
2016 Self-paced Learning for Imbalanced Data
Maciej Zieba, Jakub M. Tomczak, Jerzy Swiatek
ACIIDS (1)2
2016 Ensemble boosted trees with synthetic features generation in application to bankruptcy prediction
Maciej Zieba, Sebastian K. Tomczak, Jakub M. Tomczak
Expert Syst. Appl.3
2016 Articulated tracking with manifold regularized particle filter
abstract
In this paper, we investigate articulated human motion tracking from video sequences using Bayesian approach. We derive a generic particle-based filtering procedure with a low-dimensional manifold. The manifold can be treated as a regularizer that enforces a distribution over poses during tracking process to be concentrated around the low-dimensional embedding. We refer to our method as manifold regularized particle filter . We present a particular implementation of our method based on back-constrained gaussian process latent variable model and gaussian diffusion. The proposed approach is evaluated using the real-life benchmark dataset HumanEva . We show empirically that the presented sampling scheme outperforms sampling-importance resampling and annealed particle filter procedures.
Adam Gonczarek, Jakub M. Tomczak
Mach. Vis. Appl.2
2016 Learning Informative Features from Restricted Boltzmann Machines
abstract
In recent years deep learning paradigm achieved important empirical success in a number of practical applications such as object recognition, speech recognition and natural language processing. A lot of effort has been put on understanding theoretical aspects of this success, however, still there is no common view on how deep architectures should be trained and thus many open questions remain. One hypothesis focuses on formulating good criterion (prior) that may help to learn a set of features capable of disentangling hidden factors. Following this line of thinking, in this paper, we propose to add a penalty (regularization) term to the log-likelihood function that enforces hidden units to maximize entropy and to be pairwise uncorrelated, for given observables. We hypothesize that the proposed framework for learning informative features results in more discriminative data representation that maintains its generative capabilities. In order to verify our hypothesis we apply the regularization term to the Restricted Boltzmann Machine (RBM) and carry out empirical study on three classification problems: character recognition, object recognition, and document classification. The experiments confirm that the proposed approach indeed increases discriminative and generative performance in comparison to RBM trained without any regularization and with the weight-decay, the sparse regularization, the max-norm regularization, Dropout and Dropconnect .
Jakub M. Tomczak
Neural Process. Lett.1
2015 RBM-SMOTE: Restricted Boltzmann Machines for Synthetic Minority Oversampling Technique
Maciej Zieba, Jakub M. Tomczak, Adam Gonczarek
ACIIDS (1)2
2015 Classification Restricted Boltzmann Machine for comprehensible credit scoring model
Jakub M. Tomczak, Maciej Zieba
Expert Syst. Appl.1
2015 Probabilistic combination of classification rules and its application to medical diagnosis
abstract
Application of machine learning to medical diagnosis entails facing two major issues, namely, a necessity of learning comprehensible models and a need of coping with imbalanced data phenomenon. The first one corresponds to a problem of implementing interpretable models, e.g., classification rules or decision trees. The second issue represents a situation in which the number of examples from one class (e.g., healthy patients) is significantly higher than the number of examples from the other class (e.g., ill patients). Learning algorithms which are prone to the imbalance data return biased models towards the majority class. In this paper, we propose a probabilistic combination of soft rules , which can be seen as a probabilistic version of the classification rules, by introducing new latent random variable called conjunctive feature . The conjunctive features represent conjunctions of values of attribute variables (features) and we assume that for given conjunctive feature the object and its label (class) become independent random variables. In order to deal with the between class imbalance problem, we present a new estimator which incorporates the knowledge about data imbalanceness into hyperparameters of initial probability of objects with fixed class labels. Additionally, we propose a method for aggregating sufficient statistics needed to estimate probabilities in a graph-based structure to speed up computations. At the end, we carry out two experiments: (1) using benchmark datasets, (2) using medical datasets. The results are discussed and the conclusions are drawn.
Jakub M. Tomczak, Maciej Zieba
Mach. Learn.1
2015 Boosted SVM with active learning strategy for imbalanced data
abstract
In this work, we introduce a novel training method for constructing boosted Support Vector Machines (SVMs) directly from imbalanced data. The proposed solution incorporates the mechanisms of active learning strategy to eliminate redundant instances and more properly estimate misclassification costs for each of the base SVMs in the committee. To evaluate our approach, we make comprehensive experimental studies on the set of $$44$$ benchmark datasets with various types of imbalance ratio. In addition, we present application of our method to the real-life decision problem related to the short-term loans repayment prediction.
Maciej Zieba, Jakub M. Tomczak
Soft Comput.2
2014 Sparse hidden units activation in Restricted Boltzmann Machine
Jakub M. Tomczak, Adam Gonczarek
ICSEng1
2014 Accelerated learning for Restricted Boltzmann Machine with momentum term
Szymon Zareba, Adam Gonczarek, Jakub M. Tomczak, Jerzy Swiatek
ICSEng3
2014 Selecting right questions with Restricted Boltzmann Machines
Maciej Zieba, Jakub M. Tomczak, Krzysztof Brzostowski
ICSEng2
2013 Decision rules extraction from data stream in the presence of changing context for diabetes treatment
abstract
The knowledge extraction is an important element of the e-Health system. In this paper, we introduce a new method for decision rules extraction called Graph-based Rules Inducer to support the medical interview in the diabetes treatment. The emphasis is put on the capability of hidden context change tracking. The context is understood as a set of all factors affecting patient condition. In order to follow context changes, a forgetting mechanism with a forgetting factor is implemented in the proposed algorithm. Moreover, to aggregate data, a graph representation is used and a limitation of the search space is proposed to protect from overfitting. We demonstrate the advantages of our approach in comparison with other methods through an empirical study on the Electricity benchmark data set in the classification task. Subsequently, our method is applied in the diabetes treatment as a tool supporting medical interviews.
Jakub M. Tomczak, Adam Gonczarek
Knowl. Inf. Syst.1
2012 A Probabilistic Approach to Structural Change Prediction in Evolving Social Networks
abstract
We propose a predictive model of structural changes in elementary sub graphs of social network based on Mixture of Markov Chains. The model is trained and verified on a dataset from a large corporate social network analyzed in short, one day-long time windows, and reveals distinctive patterns of evolution of connections on the level of local network topology. We argue that the network investigated in such short timescales is highly dynamic and therefore immune to classic methods of link prediction and structural analysis, and show that in the case of complex networks, the dynamic sub graph mining may lead to better prediction accuracy. The experiments were carried out on the logs from the Wroclaw University of Technology mail server.
Krzysztof Juszczyszyn, Adam Gonczarek, Jakub M. Tomczak, Katarzyna Musial, Marcin Budka
ASONAM3
2011 Context Change Detection for Resource Allocation in Service-Oriented Systems
Piotr Rygielski, Jakub M. Tomczak
KES (2)2
2010 Student Courses Recommendation Using Ant Colony Optimization
Janusz Sobecki, Jakub M. Tomczak
ACIIDS (2)2