Tomasz Trzcinski

dblp:05/11408 · DBLP profile ↗
← Back
84ranked-venue papers
8as first author
64since 2021 · last 2026
0000-0002-1486-8906ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 63 · 7 first-author · 50 since 2021Graphics, computer vision, multimedia, augmented reality and games · 35 · 4 first-author · 27 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 7 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Theory of computation · 2Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Reducing Estimation Uncertainty Using Normalizing Flows and Stratification
Pawel Lorek, Rafal Nowak, Rafal Topolnicki, Tomasz Trzcinski, Maciej Zieba, Aleksandra Krystecka
ACIIDS (1)4
2025 ExpertSim: Fast Particle Detector Simulation Using Mixture-of-Generative-Experts
abstract
Simulating detector responses is a crucial part of understanding the inner workings of particle collisions in the Large Hadron Collider at CERN. Such simulations are currently performed with statistical Monte Carlo methods, which are computationally expensive and put a significant strain on CERN’s computational grid. Therefore, recent proposals advocate for generative machine learning methods to enable more efficient simulations. However, the distribution of the data varies significantly across the simulations, which is hard to capture with out-of-the-box methods. In this study, we present ExpertSim - a deep learning simulation approach tailored for the Zero Degree Calorimeter in the ALICE experiment. Our method utilizes a Mixture-of-Generative-Experts architecture, where each expert specializes in simulating a different subset of the data. This allows for a more precise and efficient generation process, as each expert focuses on a specific aspect of the calorimeter response. ExpertSim not only improves accuracy, but also provides a significant speedup compared to the traditional Monte-Carlo methods, offering a promising solution for high-efficiency detector simulations in particle physics experiments at CERN. We make the code available at https://github.com/patrick-bedkowski/expertsim-mix-of-generative-experts.
Patryk Bedkowski, Jan Dubinski, Filip Szatkowski, Kamil Deja, Przemyslaw Rokita, Tomasz Trzcinski
ECAI6
2025 GUIDE: Guidance-Based Incremental Learning with Diffusion Models
abstract
Deep neural networks often forget previously learned information when trained sequentially on new objectives, a phenomenon known as catastrophic forgetting. Existing generative strategies combat this issue by randomly sampling rehearsal examples from a generative model. Such an approach contradicts buffer-based approaches where sampling strategy plays an important role. We propose to bridge this gap and benefit from the combination of DDPM trained on the previous task and the classifier guidance technique to actively generate rehearsal examples specifically designed to minimize forgetting in the currently trained classifier. Our experimental results show that GUIDE significantly reduces catastrophic forgetting, outperforming conventional random sampling approaches and surpassing recent state-of-the-art methods in continual learning with generative replay, and buffer-based rehearsal.
Bartosz Cywinski, Kamil Deja, Tomasz Trzcinski, Bartlomiej Twardowski, Lukasz Kucinski
ECAI3
2025 How to Train Your Multi-Exit Model? Analyzing the Impact of Training Strategies
abstract
Early exits enable the network's forward pass to terminate early by attaching trainable internal classifiers to the backbone network. Existing early-exit methods typically adopt either a joint training approach, where the backbone and exit heads are trained simultaneously, or a disjoint approach, where the heads are trained separately. However, the implications of this choice are often overlooked, with studies typically adopting one approach without adequate justification. This choice influences training dynamics and its impact remains largely unexplored. In this paper, we introduce a set of metrics to analyze early-exit training dynamics and guide the choice of training strategy. We demonstrate that conventionally used joint and disjoint regimes yield suboptimal performance. To address these limitations, we propose a mixed training strategy: the backbone is trained first, followed by the training of the entire multi-exit network. Through comprehensive evaluations of training strategies across various architectures, datasets, and early-exit methods we present strengths and weaknesses of the early exit training strategies. In particular, we show consistent improvements in performance and efficiency using the proposed mixed strategy.
Piotr Kubaty, Bartosz Wójcik, Bartlomiej Krzepkowski, Monika Michaluk, Tomasz Trzcinski, Jary Pomponi, Kamil Adamczewski
ICML5
2025 Improving Continual Learning Performance and Efficiency with Auxiliary Classifiers
abstract
Continual learning is crucial for applying machine learning in challenging, dynamic, and often resource-constrained environments. However, catastrophic forgetting — overwriting previously learned knowledge when new information is acquired — remains a major challenge. In this work, we examine the intermediate representations in neural network layers during continual learning and find that such representations are less prone to forgetting, highlighting their potential to accelerate computation. Motivated by these findings, we propose to use auxiliary classifiers (ACs) to enhance performance and demonstrate that integrating ACs into various continual learning methods consistently improves accuracy across diverse evaluation settings, yielding an average 10% relative gain. We also leverage the ACs to reduce the average cost of the inference by 10-60% without compromising accuracy, enabling the model to return the predictions before computing all the layers. Our approach provides a scalable and efficient solution for continual learning.
Filip Szatkowski, Yaoyue Zheng, Fei Yang 0004, Tomasz Trzcinski, Bartlomiej Twardowski, Joost van de Weijer 0001
ICML4
2025 RegScore: Scoring Systems for Regression Tasks
abstract
Scoring systems are widely adopted in medical applications for their inherent simplicity and transparency, particularly for classification tasks involving tabular data. In this work, we introduce RegScore, a novel, sparse, and interpretable scoring system specifically designed for regression tasks. Unlike conventional scoring systems constrained to integer-valued coefficients, RegScore leverages beam search and k-sparse ridge regression to relax these restrictions, thus enhancing predictive performance. We extend RegScore to bimodal deep learning by integrating tabular data with medical images. We utilize the classification token from the TIP (Tabular Image Pretraining) transformer to generate Personalized Linear Regression parameters and a Personalized RegScore, enabling individualized scoring. We demonstrate the effectiveness of RegScore by estimating mean Pulmonary Artery Pressure using tabular data and further refine these estimates by incorporating cardiac MRI images. Experimental results show that RegScore and its personalized bimodal extensions achieve performance comparable to, or better than, state-of-the-art black-box models. Our method provides a transparent and interpretable approach for regression tasks in clinical settings, promoting more informed and trustworthy decision-making. We provide our code at https://github.com/SanoScience/RegScore .
Michal K. Grzeszczyk, Tomasz Szczepanski, Pawel Renc, Siyeop Yoon, Jerome Charton, Tomasz Trzcinski, Arkadiusz Sitek
MICCAI (14)6
2025 GEPAR3D: Geometry Prior-Assisted Learning for 3D Tooth Segmentation
abstract
Tooth segmentation in Cone-Beam Computed Tomography (CBCT) remains challenging, especially for fine structures like root apices, which is critical for assessing root resorption in orthodontics. We introduce GEPAR3D, a novel approach that unifies instance detection and multi-class segmentation into a single step tailored to improve root segmentation. Our method integrates a Statistical Shape Model of dentition as a geometric prior, capturing anatomical context and morphological consistency without enforcing restrictive adjacency constraints. We leverage a deep watershed method, modeling each tooth as a continuous 3D energy basin encoding voxel distances to boundaries. This instance-aware representation ensures accurate segmentation of narrow, complex root apices. Trained on publicly available CBCT scans from a single center, our method is evaluated on external test sets from two in-house and two public medical centers. GEPAR3D achieves the highest overall segmentation performance, averaging a Dice Similarity Coefficient (DSC) of 95.0% (+2.8% over the second-best method) and increasing recall to 95.2% (+9.5%) across all test sets. Qualitative analyses demonstrated substantial improvements in root segmentation quality, indicating significant potential for more accurate root resorption assessment and enhanced clinical decision-making in orthodontics. We provide the implementation and dataset at github.com/tomek1911/GEPAR3D .
Tomasz Szczepanski, Szymon Plotka, Michal K. Grzeszczyk, Arleta Adamowicz, Piotr Fudalej, Przemyslaw Korzeniowski, Tomasz Trzcinski, Arkadiusz Sitek
MICCAI (2)7
2025 1000 Layer Networks for Self-Supervised RL: Scaling Depth Can Enable New Goal-Reaching Capabilities
abstract
Scaling up self-supervised learning has driven breakthroughs in language and vision, yet comparable progress has remained elusive in reinforcement learning (RL). In this paper, we study building blocks for self-supervised RL that unlock substantial improvements in scalability, with network depth serving as a critical factor. Whereas most RL papers in recent years have relied on shallow architectures (around 2 -- 5 layers), we demonstrate that increasing the depth up to 1024 layers can significantly boost performance. Our experiments are conducted in an unsupervised goal-conditioned setting, where no demonstrations or rewards are provided, so an agent must explore (from scratch) and learn how to maximize the likelihood of reaching commanded goals. Evaluated on simulated locomotion and manipulation tasks, our approach increases performance on the self-supervised contrastive RL algorithm by $2\times$ -- $50\times$, outperforming other goal-conditioned baselines. Increasing the model depth not only increases success rates but also qualitatively changes the behaviors learned.
Ishaan Javali, Michal Bortkiewicz, Tomasz Trzcinski, Benjamin Eysenbach
NeurIPS4
2025 LumiGauss: Relightable Gaussian Splatting in the Wild
abstract
Decoupling lighting from geometry using unconstrained photo collections is notoriously challenging. Solving it would benefit many users as creating complex 3D assets takes days of manual labor. Many previous works have attempted to address this issue, often at the expense of output fidelity, which questions the practicality of such methods. We introduce LumiGauss - a technique that tackles 3D reconstruction of scenes and environmental lighting through 2D Gaussian Splatting. Our approach yields high-quality scene reconstructions and enables realistic lighting synthesis under novel environment maps. We also propose a method for enhancing the quality of shadows, common in outdoor scenes, by exploiting spherical harmonics properties. Our approach facilitates seamless integration with game engines and enables the use of fast precomputed radiance transfer. We validate our method on the NeRF-OSR dataset, demonstrating superior performance over baseline methods. Moreover, LumiGauss can synthesize realistic images for unseen environment maps. Our code: https://github.com/joaxkal/lumigauss.
Joanna Kaleta, Kacper Kania, Tomasz Trzcinski, Marek Kowalski
WACV3
2025 Exploring the Stability Gap in Continual Learning: The Role of the Classification Head
abstract
Continual learning (CL) has emerged as a critical area in machine learning, enabling neural networks to learn from evolving data distributions while mitigating catastrophic forgetting. However, recent research has identified the stability gap - a phenomenon where models initially lose performance on previously learned tasks before partially recovering during training. Such learning dynamics are contradictory to the intuitive understanding of stability in continual learning where one would expect the performance to degrade gradually instead of rapidly decreasing and then partially recovering later. To better understand and alleviate the stability gap, we investigate it at different levels of the neural network architecture, particularly focusing on the role of the classification head. We introduce the nearest-mean classifier (NMC) as a tool to attribute the influence of the backbone and the classification head on the stability gap. Our experiments demonstrate that NMC not only improves final performance, but also significantly enhances training stability across various continual learning benchmarks, including CIFAR100, ImageNet100, CUB-200, and FGVC Aircrafts. Moreover, we find that NMC also reduces task-recency bias. Our analysis provides new insights into the stability gap and suggests that the primary contributor to this phenomenon is the linear head, rather than the insufficient representation learning.
Wojciech Lapacz, Daniel Marczak, Filip Szatkowski, Tomasz Trzcinski
WACV4
2025 Beyond Grids: Exploring Elastic Input Sampling for Vision Transformers
Adam Pardyl, Grzegorz Kurzejamski, Jan Olszewski, Tomasz Trzcinski, Bartosz Zielinski 0001
WACV4
2025 Adapt & Align: Continual Learning with Generative Models' Latent Space Alignment
abstract
Motivation: Neural networks suffer from abrupt loss in performance when retrained with additional data from different distributions. At the same time, training with additional data without access to the previous examples rarely improves the model’s performance. Methods: We propose Adapt & Align, a novel continual learning framework that leverages generative models to align their latent representations across tasks. The approach is divided into two phases: • Local Training: Train a generative model (e.g., a Variational Autoencoder (VAE) or a Generative Adversarial Network (GAN)) on the current task to capture task-specific features. • Global Training: Use a translator network to map these task-specific latent representations into a unified global latent space, thereby facilitating both forward and backward knowledge transfer. Results: Experiments on benchmark datasets (e.g., MNIST, Omniglot, CIFAR, CelebA) as well as real-world application for particle simulation at CERN demonstrate that Adapt & Align mitigates catastrophic forgetting and improves generation quality as indicated by metrics such as Fréchet Inception Distance (FID), distribution precision and recall, or accuracy for the downstream classification task. Ablation studies confirm the critical role of each component.
Kamil Deja, Bartosz Cywinski, Jan Rybarczyk, Tomasz Trzcinski
Neurocomputing4
2024 HyperCube: Implicit Field Representations of Voxelized 3D Models (Student Abstract)
abstract
Implicit field representations offer an effective way of generating 3D object shapes. They leverage an implicit decoder (IM-NET) trained to take a 3D point coordinate concatenated with a shape encoding and to output a value indicating whether the point is outside the shape. This approach enables the efficient rendering of visually plausible objects but also has some significant limitations, resulting in a cumbersome training procedure and empty spaces within the rendered mesh. In this paper, we introduce a new HyperCube architecture based on interval arithmetic that enables direct processing of 3D voxels, trained using a hypernetwork paradigm to enforce model convergence. The code is available at https://github.com/mproszewska/hypercube.
Magdalena Proszewska, Marcin Mazur, Tomasz Trzcinski, Przemyslaw Spurek
AAAI3
2024 AR-TTA: A Simple Method for Real-World Continual Test-Time Adaptation
Damian Sójka, Bartlomiej Twardowski, Tomasz Trzcinski, Sebastian Cygert
BMVC3
2024 Zero-Waste Machine Learning
abstract
Today, both science and industry rely heavily on machine learning models, predominantly artificial neural networks, that become increasingly complex and demand more computing resources to be trained. In this paper, we will look holistically at the efficiency of machine learning models and draw the inspirations to address their main challenges from the green sustainable economy principles. Instead of constraining some computations or memory used by the models, we will focus on reusing what is available to them: computations done in the previous processing steps, partial information accessible at run-time, or knowledge gained by the model during previous training sessions in continually learned models. This new research path of zero-waste machine learning can lead to several research questions related to efficiency of contemporary neural networks - how machine learning models can learn better with less data? How they select relevant data samples out of many? Finally, how can they build on top of already trained models to reduce the need for more training samples? Here, we explore all the above questions and attempt to answer them.
Tomasz Trzcinski, Bartlomiej Twardowski, Bartosz Zielinski 0001, Kamil Adamczewski, Bartosz Wójcik
ECAI1
2024 Revisiting Supervision for Continual Representation Learning
Daniel Marczak, Sebastian Cygert, Tomasz Trzcinski, Bartlomiej Twardowski
ECCV (6)3
2024 MAGMAX: Leveraging Model Merging for Seamless Continual Learning
Daniel Marczak, Bartlomiej Twardowski, Tomasz Trzcinski, Sebastian Cygert
ECCV (85)3
2024 AdaGlimpse: Active Visual Exploration with Arbitrary Glimpse Position and Scale
Adam Pardyl, Michal Wronka, Maciej Wolczyk, Kamil Adamczewski, Tomasz Trzcinski, Bartosz Zielinski 0001
ECCV (21)5
2024 Category Adaptation Meets Projected Distillation in Generalized Continual Category Discovery
Grzegorz Rypesc, Daniel Marczak, Sebastian Cygert, Tomasz Trzcinski, Bartlomiej Twardowski
ECCV (11)4
2024 CLIP-DINOiser: Teaching CLIP a Few DINO Tricks for Open-Vocabulary Semantic Segmentation
Monika Wysoczanska, Oriane Siméoni, Michaël Ramamonjisoa, Andrei Bursuc, Tomasz Trzcinski, Patrick Pérez
ECCV (61)5
2024 Subgoal Reachability in Goal Conditioned Hierarchical Reinforcement Learning
Michal Bortkiewicz, Jakub Lyskawa, Pawel Wawrzynski, Mateusz Ostaszewski, Artur Grudkowski, Bartlomiej Sobieski, Tomasz Trzcinski
ICAART (1)7
2024 Divide and not forget: Ensemble of selectively trained experts in Continual Learning
abstract
Class-incremental learning is becoming more popular as it helps models widen their applicability while not forgetting what they already know. A trend in this area is to use a mixture-of-expert technique, where different models work together to solve the task. However, the experts are usually trained all at once using whole task data, which makes them all prone to forgetting and increasing computational burden. To address this limitation, we introduce a novel approach named SEED. SEED selects only one, the most optimal expert for a considered task, and uses data from this task to fine-tune only this expert. For this purpose, each expert represents each class with a Gaussian distribution, and the optimal expert is selected based on the similarity of those distributions. Consequently, SEED increases diversity and heterogeneity within the experts while maintaining the high stability of this ensemble method. The extensive experiments demonstrate that SEED achieves state-of-the-art performance in exemplar-free settings across various scenarios, showing the potential of expert diversification through data in continual learning.
Grzegorz Rypesc, Sebastian Cygert, Valeriya Khan, Tomasz Trzcinski, Bartosz Zielinski 0001, Bartlomiej Twardowski
ICLR4
2024 Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning
abstract
Recent advancements in off-policy Reinforcement Learning (RL) have significantly improved sample efficiency, primarily due to the incorporation of various forms of regularization that enable more gradient update steps than traditional agents. However, many of these techniques have been tested in limited settings, often on tasks from single simulation benchmarks and against well-known algorithms rather than a range of regularization approaches. This limits our understanding of the specific mechanisms driving RL improvements. To address this, we implemented over 60 different off-policy agents, each integrating established regularization techniques from recent state-of-the-art algorithms. We tested these agents across 14 diverse tasks from 2 simulation benchmarks, measuring training metrics related to overestimation, overfitting, and plasticity loss — issues that motivate the examined regularization techniques. Our findings reveal that while the effectiveness of a specific regularization setup varies with the task, certain combinations consistently demonstrate robust and superior performance. Notably, a simple Soft Actor-Critic agent, appropriately regularized, reliably finds a better-performing policy within the training regime, which previously was achieved mainly through model-based approaches.
Michal Nauman, Michal Bortkiewicz, Piotr Milos, Tomasz Trzcinski, Mateusz Ostaszewski, Marek Cygan
ICML4
2024 TabMixer: Noninvasive Estimation of the Mean Pulmonary Artery Pressure via Imaging and Tabular Data Mixing
abstract
Right Heart Catheterization is a gold standard procedure for diagnosing Pulmonary Hypertension by measuring mean Pulmonary Artery Pressure (mPAP). It is invasive, costly, time-consuming and carries risks. In this paper, for the first time, we explore the estimation of mPAP from videos of noninvasive Cardiac Magnetic Resonance Imaging. To enhance the predictive capabilities of Deep Learning models used for this task, we introduce an additional modality in the form of demographic features and clinical measurements. Inspired by all-Multilayer Perceptron architectures, we present TabMixer, a novel module enabling the integration of imaging and tabular data through spatial, temporal and channel mixing. Specifically, we present the first approach that utilizes Multilayer Perceptrons to interchange tabular information with imaging features in vision models. We test TabMixer for mPAP estimation and show that it enhances the performance of Convolutional Neural Networks, 3D-MLP and Vision Transformers while being competitive with previous modules for imaging and tabular data. Our approach has the potential to improve clinical processes involving both modalities, particularly in noninvasive mPAP estimation, thus, significantly enhancing the quality of life for individuals affected by Pulmonary Hypertension. We provide a source code for using TabMixer at https://github.com/SanoScience/TabMixer .
Michal K. Grzeszczyk, Przemyslaw Korzeniowski, Samer Alabed, Andrew J. Swift, Tomasz Trzcinski, Arkadiusz Sitek
MICCAI (5)5
2024 Let Me DeCode You: Decoder Conditioning with Tabular Data
abstract
Training deep neural networks for 3D segmentation tasks can be challenging, often requiring efficient and effective strategies to improve model performance. In this study, we introduce a novel approach, DeCode, that utilizes label-derived features for model conditioning to support the decoder in the reconstruction process dynamically, aiming to enhance the efficiency of the training process. DeCode focuses on improving 3D segmentation performance through the incorporation of conditioning embedding with learned numerical representation of 3D-label shape features. Specifically, we develop an approach, where conditioning is applied during the training phase to guide the network toward robust segmentation. When labels are not available during inference, our model infers the necessary conditioning embedding directly from the input data, thanks to a feed-forward network learned during the training phase. This approach is tested using synthetic data and cone-beam computed tomography (CBCT) images of teeth. For CBCT, three datasets are used: one publicly available and two in-house. Our results show that DeCode significantly outperforms traditional, unconditioned models in terms of generalization to unseen data, achieving higher accuracy at a reduced computational cost. This work represents the first of its kind to explore conditioning strategies in 3D data segmentation, offering a novel and more efficient method for leveraging annotated data. Our code, pre-trained models are publicly available at https://github.com/SanoScience/DeCode .
Tomasz Szczepanski, Michal K. Grzeszczyk, Szymon Plotka, Arleta Adamowicz, Piotr Fudalej, Przemyslaw Korzeniowski, Tomasz Trzcinski, Arkadiusz Sitek
MICCAI (3)7
2024 Task-recency bias strikes back: Adapting covariances in Exemplar-Free Class Incremental Learning
abstract
Exemplar-Free Class Incremental Learning (EFCIL) tackles the problem of training a model on a sequence of tasks without access to past data. Existing state-of-the-art methods represent classes as Gaussian distributions in the feature extractor's latent space, enabling Bayes classification or training the classifier by replaying pseudo features. However, we identify two critical issues that compromise their efficacy when the feature extractor is updated on incremental tasks. First, they do not consider that classes' covariance matrices change and must be adapted after each task. Second, they are susceptible to a task-recency bias caused by dimensionality collapse occurring during training. In this work, we propose AdaGauss - a novel method that adapts covariance matrices from task to task and mitigates the task-recency bias owing to the additional anti-collapse loss function. AdaGauss yields state-of-the-art results on popular EFCIL benchmarks and datasets when training from scratch or starting from a pre-trained backbone.
Grzegorz Rypesc, Sebastian Cygert, Tomasz Trzcinski, Bartlomiej Twardowski
NeurIPS3
2024 MISS: Multiclass Interpretable Scoring Systems
abstract
In this work, we present a novel, machine-learning approach for constructing Multiclass Interpretable Scoring Systems (MISS) - a fully data-driven methodology for generating single, sparse, and user-friendly scoring systems for multi-class classification problems. Scoring systems are commonly utilized as decision support models in healthcare, criminal justice, and other domains where interpretability of predictions and ease of use are crucial. Prior methods for data-driven scoring, such as SLIM (Supersparse Linear Integer Model), were limited to binary classification tasks and extensions to multiclass domains were primarily accomplished via one-versus-all-type techniques. The scores produced by our method can be easily transformed into class probabilities via the softmax function. We demonstrate techniques for dimensionality reduction and heuristics that enhance the training efficiency and decrease the optimality gap, a measure that can certify the optimality of the model. Our approach has been extensively evaluated on datasets from various domains, and the results indicate that it is competitive with other machine learning models in terms of classification performance metrics and provides well-calibrated class probabilities.
Michal K. Grzeszczyk, Tomasz Trzcinski, Arkadiusz Sitek
SDM2
2024 Towards More Realistic Membership Inference Attacks on Large Diffusion Models
abstract
Generative diffusion models, including Stable Diffusion and Midjourney, can generate visually appealing, diverse, and high-resolution images for various applications. These models are trained on billions of internet-sourced images, raising significant concerns about the potential unauthorized use of copyright-protected images. In this paper, we examine whether it is possible to determine if a specific image was used in the training set, a problem known in the cybersecurity community as a membership inference attack. Our focus is on Stable Diffusion, and we address the challenge of designing a fair evaluation framework to answer this membership question. We propose a new dataset to establish a fair evaluation setup and apply it to Stable Diffusion, also applicable to other generative models. With the proposed dataset, we execute membership attacks (both known and newly introduced). Our research reveals that previously proposed evaluation setups do not provide a full understanding of the effectiveness of membership inference attacks. We conclude that the membership inference attack remains a significant challenge for large diffusion models (often deployed as black-box systems), indicating that related privacy and copyright issues will persist in the foreseeable future.
Jan Dubinski, Antoni Kowalczuk, Stanislaw Pawlak, Przemyslaw Rokita, Tomasz Trzcinski, Pawel Morawiecki
WACV5
2024 Adapt Your Teacher: Improving Knowledge Distillation for Exemplar-free Continual Learning
abstract
In this work, we investigate exemplar-free class incremental learning (CIL) with knowledge distillation (KD) as a regularization strategy, aiming to prevent forgetting. KD-based methods are successfully used in CIL, but they often struggle to regularize the model without access to exemplars of the training data from previous tasks. Our analysis reveals that this issue originates from substantial representation shifts in the teacher network when dealing with out-of-distribution data. This causes large errors in the KD loss component, leading to performance degradation in CIL models. Inspired by recent test-time adaptation methods, we introduce Teacher Adaptation (TA), a method that concurrently updates the teacher and the main models during incremental training. Our method seamlessly integrates with KD-based CIL approaches and allows for consistent enhancement of their performance across multiple exemplar-free CIL benchmarks. The source code for our method is available at https://github.com/fszatkowski/cl-teacher-adaptation.
Filip Szatkowski, Mateusz Pyla, Marcin Przewiezlikowski, Sebastian Cygert, Bartlomiej Twardowski, Tomasz Trzcinski
WACV6
2024 CLIP-DIY: CLIP Dense Inference Yields Open-Vocabulary Semantic Segmentation For-Free
abstract
The emergence of CLIP has opened the way for open-world image perception. The zero-shot classification capabilities of the model are impressive but are harder to use for dense tasks such as image segmentation. Several methods have proposed different modifications and learning schemes to produce dense output. Instead, we propose in this work an open-vocabulary semantic segmentation method, dubbed CLIP-DIY, which does not require any additional training or annotations, but instead leverages existing unsupervised object localization approaches. In particular, CLIP-DIY is a multi-scale approach that directly exploits CLIP classification abilities on patches of different sizes and aggregates the decision in a single map. We further guide the segmentation using foreground/background scores obtained using unsupervised object localization methods. With our method, we obtain state-of-the-art zero-shot semantic segmentation results on PASCAL VOC and perform on par with the best methods on COCO.
Monika Wysoczanska, Michaël Ramamonjisoa, Tomasz Trzcinski, Oriane Siméoni
WACV3
2024 Efficient GPU implementation of randomized SVD and its applications
Lukasz Struski, Pawel M. Morkisz, Przemyslaw Spurek, Samuel Rodriguez Bernabeu, Tomasz Trzcinski
Expert Syst. Appl.5
2024 Points2NeRF: Generating Neural Radiance Fields from 3D point cloud
Dominik Zimny, Joanna Waczynska, Tomasz Trzcinski, Przemyslaw Spurek
Pattern Recognit. Lett.3
2023 BlendFields: Few-Shot Example-Driven Facial Modeling
abstract
Generating faithful visualizations of human faces requires capturing both coarse and fine-level details of the face geometry and appearance. Existing methods are either data-driven, requiring an extensive corpus of data not publicly accessible to the research community, or fail to capture fine details because they rely on geometric face models that cannot represent fine-grained details in texture with a mesh discretization and linear deformation designed to model only a coarse face geometry. We introduce a method that bridges this gap by drawing inspiration from traditional computer graphics techniques. Unseen expressions are modeled by blending appearance from a sparse set of extreme poses. This blending is performed by measuring local volumetric changes in those expressions and locally reproducing their appearance whenever a similar expression is performed at test time. We show that our method generalizes to unseen expressions, adding fine-grained effects on top of smooth volumetric deformations of a face, and demonstrate how it generalizes beyond faces.
Kacper Kania, Stephan J. Garbin, Andrea Tagliasacchi, Virginia Estellers, Kwang Moo Yi, Julien Valentin, Tomasz Trzcinski, Marek Kowalski
CVPR7
2023 Active Visual Exploration Based on Attention-Map Entropy
abstract
Active visual exploration addresses the issue of limited sensor capabilities in real-world scenarios, where successive observations are actively chosen based on the environment. To tackle this problem, we introduce a new technique called Attention-Map Entropy (AME). It leverages the internal uncertainty of the transformer-based model to determine the most informative observations. In contrast to existing solutions, it does not require additional loss components, which simplifies the training. Through experiments, which also mimic retina-like sensors, we show that such simplified training significantly improves the performance of reconstruction, segmentation and classification on publicly available datasets.
Adam Pardyl, Grzegorz Rypesc, Grzegorz Kurzejamski, Bartosz Zielinski 0001, Tomasz Trzcinski
IJCAI5
2023 TabAttention: Learning Attention Conditionally on Tabular Data
Michal K. Grzeszczyk, Szymon Plotka, Beata Rebizant, Katarzyna Kosinska-Kaczynska, Michal Lipa, Robert Brawura-Biskupski-Samaha, Przemyslaw Korzeniowski, Tomasz Trzcinski, Arkadiusz Sitek
MICCAI (7)8
2023 Bucks for Buckets (B4B): Active Defenses Against Stealing Encoders
abstract
Machine Learning as a Service (MLaaS) APIs provide ready-to-use and high-utility encoders that generate vector representations for given inputs. Since these encoders are very costly to train, they become lucrative targets for model stealing attacks during which an adversary leverages query access to the API to replicate the encoder locally at a fraction of the original training costs. We propose *Bucks for Buckets (B4B)*, the first *active defense* that prevents stealing while the attack is happening without degrading representation quality for legitimate API users. Our defense relies on the observation that the representations returned to adversaries who try to steal the encoder's functionality cover a significantly larger fraction of the embedding space than representations of legitimate users who utilize the encoder to solve a particular downstream task. B4B leverages this to adaptively adjust the utility of the returned representations according to a user's coverage of the embedding space. To prevent adaptive adversaries from eluding our defense by simply creating multiple user accounts (sybils), B4B also individually transforms each user's representations. This prevents the adversary from directly aggregating representations over multiple accounts to create their stolen encoder copy. Our active defense opens a new path towards securely sharing and democratizing encoders over public APIs.
Jan Dubinski, Stanislaw Pawlak, Franziska Boenisch, Tomasz Trzcinski, Adam Dziedzic
NeurIPS4
2023 The Tunnel Effect: Building Data Representations in Deep Neural Networks
abstract
Deep neural networks are widely known for their remarkable effectiveness across various tasks, with the consensus that deeper networks implicitly learn more complex data representations. This paper shows that sufficiently deep networks trained for supervised image classification split into two distinct parts that contribute to the resulting data representations differently. The initial layers create linearly-separable representations, while the subsequent layers, which we refer to as \textit{the tunnel}, compress these representations and have a minimal impact on the overall performance. We explore the tunnel's behavior through comprehensive empirical studies, highlighting that it emerges early in the training process. Its depth depends on the relation between the network's capacity and task complexity. Furthermore, we show that the tunnel degrades out-of-distribution generalization and discuss its implications for continual learning.
Wojciech Masarczyk, Mateusz Ostaszewski, Ehsan Imani, Razvan Pascanu, Piotr Milos, Tomasz Trzcinski
NeurIPS6
2023 Learning Data Representations with Joint Diffusion Models
Kamil Deja, Tomasz Trzcinski, Jakub M. Tomczak
ECML/PKDD (2)2
2023 Hypernetworks Build Implicit Neural Representations of Sounds
Filip Szatkowski, Karol J. Piczak, Przemyslaw Spurek, Jacek Tabor, Tomasz Trzcinski
ECML/PKDD (4)5
2023 Continual learning on 3D point clouds with random compressed rehearsal
Maciej Zamorski, Michal Stypulkowski, Konrad Karanowski, Tomasz Trzcinski, Maciej Zieba
Comput. Vis. Image Underst.4
2023 Zero time waste in pre-trained early exit neural networks
abstract
The problem of reducing processing time of large deep learning models is a fundamental challenge in many real-world applications. Early exit methods strive towards this goal by attaching additional Internal Classifiers (ICs) to intermediate layers of a neural network. ICs can quickly return predictions for easy examples and, as a result, reduce the average inference time of the whole model. However, if a particular IC does not decide to return an answer early, its predictions are discarded, with its computations effectively being wasted. To solve this issue, we introduce Zero Time Waste (ZTW), a novel approach in which each IC reuses predictions returned by its predecessors by (1) adding direct connections between ICs and (2) combining previous outputs in an ensemble-like manner. We conduct extensive experiments across various multiple modes, datasets, and architectures to demonstrate that ZTW achieves a significantly better accuracy vs. inference time trade-off than other early exit methods. On the ImageNet dataset, it obtains superior results over the best baseline method in 11 out of 16 cases, reaching up to 5 percentage points of improvement on low computational budgets.
Bartosz Wójcik, Marcin Przewiezlikowski, Filip Szatkowski, Maciej Wolczyk, Klaudia Balazy, Bartlomiej Krzepkowski, Igor T. Podolak, Jacek Tabor, Marek Smieja, Tomasz Trzcinski
Neural Networks10
2023 HyperColor: A HyperNetwork Approach for Synthesizing Autocolored 3-D Models for Game Scenes Population
abstract
Designing a 3-D game scene is a tedious task that often requires a substantial amount of work. Typically, this task involves the synthesis and coloring of 3-D models within the scene. To lessen this workload, we can apply machine learning to automate some aspects of the game scene development. Earlier research has already tackled automated generation of the game scene background with machine learning. However, model autocoloring remains an underexplored problem. The automatic coloring of a 3-D model is a challenging task, especially when dealing with the digital representation of a colorful, multipart object. In such a case, we have to “understand” the object's composition and coloring scheme of each part. Moreover, existing single-stage methods have their caveats. We address these limitations by proposing a two-stage training approach to synthesize autocolored 3-D models. In the first stage, we obtain a 3-D point cloud representing a 3-D object, while in the second stage, we assign colors to points within such a cloud. Next, we generate a 3-D mesh in which the surfaces are colored based on the interpolation of colored points representing vertices of a given mesh triangle. This approach allows us to develop a smooth coloring scheme.
Ivan Kostiuk, Przemyslaw Stachura, Slawomir Konrad Tadeja, Tomasz Trzcinski, Przemyslaw Spurek
IEEE Trans. Games4
2022 Deep Learning for Automatic Detection of Qualitative Features of Lecturing
Anna Wróblewska, Jozef Jasek, Bogdan Jastrzebski, Stanislaw Pawlak, Anna Grzywacz, Siew Ann Cheong, Seng Chee Tan, Tomasz Trzcinski, Janusz A. Holyst
AIED (1)8
2022 CoNeRF: Controllable Neural Radiance Fields
abstract
We extend neural 3D representations to allow for intu-itive and interpretable user control beyond novel view ren-dering (i. e. camera control). We allow the user to annotate which part of the scene one wishes to control with just a small number of mask annotations in the training images. Our key idea is to treat the attributes as latent variables that are regressed by the neural network given the scene en-coding. This leads to afew-shot learning framework, where attributes are discovered automatically by the framework, when annotations are not provided. We apply our method to various scenes with different types of controllable attributes (e.g. expression control on human faces, or state control in movement of inanimate objects). Overall, we demonstrate, to the best of our knowledge, for the first time novel view and novel attribute re-rendering of scenes from a single video.
Kacper Kania, Kwang Moo Yi, Marek Kowalski, Tomasz Trzcinski, Andrea Tagliasacchi
CVPR4
2022 Continual Learning with Guarantees via Weight Interval Constraints
abstract
We introduce a new training paradigm that enforces interval constraints on neural network parameter space to control forgetting. Contemporary Continual Learning (CL) methods focus on training neural networks efficiently from a stream of data, while reducing the negative impact of catastrophic forgetting, yet they do not provide any firm guarantees that network performance will not deteriorate uncontrollably over time. In this work, we show how to put bounds on forgetting by reformulating continual learning of a model as a continual contraction of its parameter space. To that end, we propose Hyperrectangle Training, a new training methodology where each task is represented by a hyperrectangle in the parameter space, fully contained in the hyperrectangles of the previous tasks. This formulation reduces the NP-hard CL problem back to polynomial time while providing full resilience against forgetting. We validate our claim by developing InterContiNet (Interval Continual Learning) algorithm which leverages interval arithmetic to effectively model parameter regions as hyperrectangles. Through experimental results, we show that our approach performs well in a continual learning setup without storing data from previous tasks.
Maciej Wolczyk, Karol J. Piczak, Bartosz Wójcik, Lukasz Pustelnik, Pawel Morawiecki, Jacek Tabor, Tomasz Trzcinski, Przemyslaw Spurek
ICML7
2022 Selectively Increasing the Diversity of GAN-Generated Samples
Jan Dubinski, Kamil Deja, Sandro Wenzel, Przemyslaw Rokita, Tomasz Trzcinski
ICONIP (1)5
2022 Progressive Latent Replay for Efficient Generative Rehearsal
Stanislaw Pawlak, Filip Szatkowski, Michal Bortkiewicz, Jan Dubinski, Tomasz Trzcinski
ICONIP (4)5
2022 Multiband VAE: Latent Space Alignment for Knowledge Consolidation in Continual Learning
abstract
We propose a new method for unsupervised generative continual learning through realignment of Variational Autoencoder's latent space. Deep generative models suffer from catastrophic forgetting in the same way as other neural structures. Recent generative continual learning works approach this problem and try to learn from new data without forgetting previous knowledge. However, those methods usually focus on artificial scenarios where examples share almost no similarity between subsequent portions of data - an assumption not realistic in the real-life applications of continual learning. In this work, we identify this limitation and posit the goal of generative continual learning as a knowledge accumulation task. We solve it by continuously aligning latent representations of new data that we call bands in additional latent space where examples are encoded independently of their source task. In addition, we introduce a method for controlled forgetting of past data that simplifies this process. On top of the standard continual learning benchmarks, we propose a novel challenging knowledge consolidation scenario and show that the proposed approach outperforms state-of-the-art by up to twofold across all experiments and additional real-life evaluation. To our knowledge, Multiband VAE is the first method to show forward and backward knowledge transfer in generative continual learning.
Kamil Deja, Pawel Wawrzynski, Wojciech Masarczyk, Daniel Marczak, Tomasz Trzcinski
IJCAI5
2022 HyperPocket: Generative Point Cloud Completion
abstract
Scanning real-life scenes with modern registration devices typically give incomplete point cloud representations, mostly due to the limitations of the scanning process and 3D occlusions. Therefore, completing such partial representations remains a fundamental challenge of many computer vision applications. Most of the existing approaches aim to solve this problem by learning to reconstruct individual 3D objects in a synthetic setup of an uncluttered environment, which is far from a real-life scenario. In this work, we reformulate the problem of point cloud completion into an objects hallucination task. Thus, we introduce a novel autoencoder-based architecture called HyperPocket that disentangles latent representations and, as a result, enables the generation of multiple variants of the completed 3D point clouds. Furthermore, we split point cloud processing into two disjoint data streams and leverage a hypernetwork paradigm to fill the spaces, dubbed pockets, that are left by the missing object parts. As a result, the generated point clouds are smooth, plausible, and geometrically consistent with the scene. Moreover, our method offers competitive performances to the other state-of-the-art models, enabling a plethora of novel applications.
Przemyslaw Spurek, Artur Kasymov, Marcin Mazur, Diana Janik, Slawomir Konrad Tadeja, Lukasz Struski, Jacek Tabor, Tomasz Trzcinski
IROS8
2022 BabyNet: Residual Transformer Module for Birth Weight Prediction on Fetal Ultrasound Video
abstract
Predicting fetal weight at birth is an important aspect of perinatal care, particularly in the context of antenatal management, which includes the planned timing and the mode of delivery. Accurate prediction of weight using prenatal ultrasound is challenging as it requires images of specific fetal body parts during advanced pregnancy which is difficult to capture due to poor quality of images caused by the lack of amniotic fluid. As a consequence, predictions which rely on standard methods often suffer from significant errors. In this paper we propose the Residual Transformer Module which extends a 3D ResNet-based network for analysis of 2D+t spatio-temporal ultrasound video scans. Our end-to-end method, called BabyNet, automatically predicts fetal birth weight based on fetal ultrasound video scans. We evaluate BabyNet using a dedicated clinical set comprising 225 2D fetal ultrasound videos of pregnancies from 75 patients performed one day prior to delivery. Experimental results show that BabyNet outperforms several state-of-the-art methods and estimates the weight at birth with accuracy comparable to human experts. Furthermore, combining estimates provided by human experts with those computed by BabyNet yields the best results, outperforming either of other methods by a significant margin. The source code of BabyNet is available at https://github.com/SanoScience/BabyNet.
Szymon Plotka, Michal K. Grzeszczyk, Robert Brawura-Biskupski-Samaha, Pawel Gutaj, Michal Lipa, Tomasz Trzcinski, Arkadiusz Sitek
MICCAI (4)6
2022 On Analyzing Generative and Denoising Capabilities of Diffusion-based Deep Generative Models
abstract
Diffusion-based Deep Generative Models (DDGMs) offer state-of-the-art performance in generative modeling. Their main strength comes from their unique setup in which a model (the backward diffusion process) is trained to reverse the forward diffusion process, which gradually adds noise to the input signal. Although DDGMs are well studied, it is still unclear how the small amount of noise is transformed during the backward diffusion process. Here, we focus on analyzing this problem to gain more insight into the behavior of DDGMs and their denoising and generative capabilities. We observe a fluid transition point that changes the functionality of the backward diffusion process from generating a (corrupted) image from noise to denoising the corrupted image to the final sample. Based on this observation, we postulate to divide a DDGM into two parts: a denoiser and a generator. The denoiser could be parameterized by a denoising auto-encoder, while the generator is a diffusion-based model with its own set of parameters. We experimentally validate our proposition, showing its pros and cons.
Kamil Deja, Anna Kuzina, Tomasz Trzcinski, Jakub M. Tomczak
NeurIPS3
2022 FlowHMM: Flow-based continuous hidden Markov models
abstract
Continuous hidden Markov models (HMMs) assume that observations are generated from a mixture of Gaussian densities, limiting their ability to model more complex distributions. In this work, we address this shortcoming and propose novel continuous HMM models, dubbed FlowHMMs, that enable learning general continuous observation densities without constraining them to follow a Gaussian distribution or their mixtures. To that end, we leverage deep flow-based architectures that model complex, non-Gaussian functions and propose two variants of training a FlowHMM model. The first one, based on gradient-based technique, can be applied directly to continuous multidimensional data, yet its application to larger data sequences remains computationally expensive. Therefore, we also present a second approach to training our FlowHMM that relies on the co-occurrence matrix of discretized observations and considers the joint distribution of pairs of co-observed values, hence rendering the training time independent of the training sequence length. As a result, we obtain a model that can be flexibly adapted to the characteristics and dimensionality of the data. We perform a variety of experiments in which we compare both training strategies with a baseline of Gaussian mixture models. We show, that in terms of quality of the recovered probability distribution, accuracy of prediction of hidden states, and likelihood of unseen data, our approach outperforms the standard Gaussian methods.
Pawel Lorek, Rafal Nowak, Tomasz Trzcinski, Maciej Zieba
NeurIPS3
2022 General Hypernetwork Framework for Creating 3D Point Clouds
abstract
In this work, we propose a novel method for generating 3D point clouds that leverages the properties of hypernetworks. Contrary to the existing methods that learn only the representation of a 3D object, our approach simultaneously finds a representation of the object and its 3D surface. The main idea of our HyperCloud method is to build a hypernetwork that returns weights of a particular neural network (target network) trained to map points from prior distribution into a 3D shape. As a consequence, a particular 3D shape can be generated using point-by-point sampling from the prior distribution and transforming the sampled points with the target network. Since the hypernetwork is based on an auto-encoder architecture trained to reconstruct realistic 3D shapes, the target network weights can be considered to be a parametrization of the surface of a 3D shape, and not a standard representation of point cloud usually returned by competitive approaches. We also show that relying on hypernetworks to build 3D point cloud representations offers an elegant and flexible framework. To that point, we further extend our method by incorporating flow-based models, which results in a novel HyperFlow approach.
Przemyslaw Spurek, Maciej Zieba, Jacek Tabor, Tomasz Trzcinski
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 Large-Scale Topological Radar Localization Using Learned Descriptors
Jacek Komorowski, Monika Wysoczanska, Tomasz Trzcinski
ICONIP (2)3
2021 On Robustness of Generative Representations Against Catastrophic Forgetting
abstract
Catastrophic forgetting of previously learned knowledge while learning new tasks is a widely observed limitation of contemporary neural networks. Although many continual learning methods are proposed to mitigate this drawback, the main question remains unanswered: what is the root cause of catastrophic forgetting? In this work, we aim at answering this question by posing and validating a set of research hypotheses related to the specificity of representations built internally by neural models. More specifically, we design a set of empirical evaluations that compare the robustness of representations in discriminative and generative models against catastrophic forgetting. We observe that representations learned by discriminative models are more prone to catastrophic forgetting than their generative counterparts, which sheds new light on the advantages of developing generative models for continual learning. Finally, our work opens new research pathways and possibilities to adopt generative models in continual learning beyond mere replay mechanisms.
Wojciech Masarczyk, Kamil Deja, Tomasz Trzcinski
ICONIP (6)3
2021 FetalNet: Multi-task Deep Learning Framework for Fetal Ultrasound Biometric Measurements
Szymon Plotka, Tomasz Wlodarczyk, Adam Klasa, Michal Lipa, Arkadiusz Sitek, Tomasz Trzcinski
ICONIP (6)6
2021 Continual Learning of 3D Point Cloud Generators
abstract
Most continual learning evaluations to date have focused on fully supervised image classification problems. This work for the first time extends such an analysis to the domain of 3D point cloud generation, showing that 3D object generators are prone to catastrophic forgetting along the same vein as image classifiers. Classic mitigation techniques, such as regularization and replay, are only partially effective in alleviating this issue. We show that due to the specifics of generative tasks, it is possible to maintain most of the generative diversity with a simple technique of uniformly sampling from different columns of a progressive neural network. While such an approach performs well on a typical synthetic class-incremental setup, more realistic scenarios might hinder strong concept separation by shifting task boundaries and introducing class overlap between tasks. Therefore, we propose an autonomous branch construction (ABC) method. This learning adaptation relevant to parameter-isolation methods employs the reconstruction loss to map new training examples to proper branches of the model. Internal routing of training data allows for a more effective and robust continual learning and generation of separate concepts in overlapping task setups.
Michal Sadowski, Karol J. Piczak, Przemyslaw Spurek, Tomasz Trzcinski
ICONIP (1)4
2021 SynthTriplet GAN: Synthetic Query Expansion for Multimodal Retrieval
Ivona Tautkute, Tomasz Trzcinski
ICONIP (4)2
2021 Explaining Self-Supervised Image Representations with Visual Probing
abstract
Recently introduced self-supervised methods for image representation learning provide on par or superior results to their fully supervised competitors, yet the corresponding efforts to explain the self-supervised approaches lag behind. Motivated by this observation, we introduce a novel visual probing framework for explaining the self-supervised models by leveraging probing tasks employed previously in natural language processing. The probing tasks require knowledge about semantic relationships between image parts. Hence, we propose a systematic approach to obtain analogs of natural language in vision, such as visual words, context, and taxonomy. We show the effectiveness and applicability of those analogs in the context of explaining self-supervised representations. Our key findings emphasize that relations between language and vision can serve as an effective yet intuitive tool for discovering how machine learning models work, independently of data modality. Our work opens a plethora of research pathways towards more explainable and transparent AI.
Dominika Basaj, Witold Oleszkiewicz, Igor Sieradzki, Michal Górszczak, Barbara Rychalska, Tomasz Trzcinski, Bartosz Zielinski 0001
IJCAI6
2021 BinPlay: A Binary Latent Autoencoder for Generative Replay Continual Learning
abstract
We introduce a novel binary latent space autoen-coder architecture to rehearse training samples for the continual learning of neural networks. The ability to extend the knowledge of a model with new data without forgetting previously learned samples is a fundamental requirement in continual learning. Existing solutions address it by regularizing network weights, adjusting its architecture, or retraining with past data samples, regenerated from memory or reconstructed with generative models. Unfortunately, recreating past data from memory requires an infinite buffer, while the reconstructions of generative models tend to miss details of individual samples when generalizing beyond the training set. In this paper, we aim to overcome these limitations and introduce a novel generative rehearsal approach called BinPlay. Its main objective is to find a quality-preserving encoding of past samples into precomputed binary codes living in the autoencoder's binary latent space. Since we parametrize the formula for precomputing the codes only on the training samples' chronological indices, the autoencoder is able to compute the binary codes of rehearsed samples on the fly without the need to keep them in memory. Evaluation on three benchmark datasets shows up to a twofold accuracy improvement of BinPlay versus competing generative replay methods.
Kamil Deja, Pawel Wawrzynski, Daniel Marczak, Wojciech Masarczyk, Tomasz Trzcinski
IJCNN5
2021 MinkLoc++: Lidar and Monocular Image Fusion for Place Recognition
abstract
We introduce a discriminative multimodal descriptor based on a pair of sensor readings: a point cloud from a LiDAR and an image from an RGB camera. Our descriptor, named MinkLoc++, can be used for place recognition, re-localization and loop closure purposes in robotics or autonomous vehicles applications. We use late fusion approach, where each modality is processed separately and fused in the final part of the processing pipeline. The proposed method achieves state-of-the-art performance on standard place recognition benchmarks. We also identify dominating modality problem when training a multimodal descriptor. The problem manifests itself when the network focuses on a modality with a larger overfit to the training data. This drives the loss down during the training but leads to suboptimal performance on the evaluation set. In this work we describe how to detect and mitigate such risk when using a deep metric learning approach to train a multimodal neural network. Our code is publicly available on the project website11https://github.com/jac99/MinkLocMultimodal.
Jacek Komorowski, Monika Wysoczanska, Tomasz Trzcinski
IJCNN3
2021 Non-Gaussian Gaussian Processes for Few-Shot Regression
abstract
Gaussian Processes (GPs) have been widely used in machine learning to model distributions over functions, with applications including multi-modal regression, time-series prediction, and few-shot learning. GPs are particularly useful in the last application since they rely on Normal distributions and enable closed-form computation of the posterior probability function. Unfortunately, because the resulting posterior is not flexible enough to capture complex distributions, GPs assume high similarity between subsequent tasks - a requirement rarely met in real-world conditions. In this work, we address this limitation by leveraging the flexibility of Normalizing Flows to modulate the posterior predictive distribution of the GP. This makes the GP posterior locally non-Gaussian, therefore we name our method Non-Gaussian Gaussian Processes (NGGPs). More precisely, we propose an invertible ODE-based mapping that operates on each component of the random variable vectors and shares the parameters across all of them. We empirically tested the flexibility of NGGPs on various few-shot learning regression datasets, showing that the mapping can incorporate context embedding information to model different noise levels for periodic functions. As a result, our method shares the structure of the problem between subsequent tasks, but the contextualization allows for adaptation to dissimilarities. NGGPs outperform the competing state-of-the-art approaches on a diversified set of benchmarks and applications.
Marcin Sendera, Jacek Tabor, Aleksandra Nowak 0001, Andrzej Bedychaj, Massimiliano Patacchiola, Tomasz Trzcinski, Przemyslaw Spurek, Maciej Zieba
NeurIPS6
2021 Zero Time Waste: Recycling Predictions in Early Exit Neural Networks
abstract
The problem of reducing processing time of large deep learning models is a fundamental challenge in many real-world applications. Early exit methods strive towards this goal by attaching additional Internal Classifiers (ICs) to intermediate layers of a neural network. ICs can quickly return predictions for easy examples and, as a result, reduce the average inference time of the whole model. However, if a particular IC does not decide to return an answer early, its predictions are discarded, with its computations effectively being wasted. To solve this issue, we introduce Zero Time Waste (ZTW), a novel approach in which each IC reuses predictions returned by its predecessors by (1) adding direct connections between ICs and (2) combining previous outputs in an ensemble-like manner. We conduct extensive experiments across various datasets and architectures to demonstrate that ZTW achieves a significantly better accuracy vs. inference time trade-off than other recently proposed early exit methods.
Maciej Wolczyk, Bartosz Wójcik, Klaudia Balazy, Igor T. Podolak, Jacek Tabor, Marek Smieja, Tomasz Trzcinski
NeurIPS7
2021 Representing point clouds with generative conditional invertible flow networks
Michal Stypulkowski, Kacper Kania, Maciej Zamorski, Maciej Zieba, Tomasz Trzcinski, Jan Chorowski
Pattern Recognit. Lett.5
2020 SuperNCN: Neighbourhood Consensus Network for Robust Outdoor Scenes Matching
Grzegorz Kurzejamski, Jacek Komorowski, Lukasz Dabala, Konrad Czarnota, Simon Lynen, Tomasz Trzcinski
ACIVS6
2020 Hypernetwork approach to generating point clouds
abstract
In this work, we propose a novel method for generating 3D point clouds that leverage properties of hyper networks. Contrary to the existing methods that learn only the representation of a 3D object, our approach simultaneously finds a representation of the object and its 3D surfaces. The main idea of our HyperCloud method is to build a hyper network that returns weights of a particular neural network (target network) trained to map points from a uniform unit ball distribution into a 3D shape. As a consequence, a particular 3D shape can be generated using point-by-point sampling from the assumed prior distribution and transforming sampled points with the target network. Since the hyper network is based on an auto-encoder architecture trained to reconstruct realistic 3D shapes, the target network weights can be considered a parametrisation of the surface of a 3D shape, and not a standard representation of point cloud usually returned by competitive approaches. The proposed architecture allows to find mesh-based representation of 3D objects in a generative manner, while providing point clouds en pair in quality with the state-of-the-art methods.
Przemyslaw Spurek, Sebastian Winczowski, Jacek Tabor, Maciej Zamorski, Maciej Zieba, Tomasz Trzcinski
ICML6
2020 Plugin Networks for Inference under Partial Evidence
abstract
In this paper, we propose a novel method to incorporate partial evidence in the inference of deep convolutional neural networks. Contrary to the existing, top-performing methods, which either iteratively modify the input of the network or exploit external label taxonomy to take the partial evidence into account, we add separate network modules ("Plugin Networks") to the intermediate layers of a pretrained convolutional network. The goal of these modules is to incorporate additional signal, i.e. information about known labels, into the inference procedure, and adjust the predicted output accordingly. Since the attached plugins have a simple structure, consisting of only fully connected layers, we drastically reduced the computational cost of training and inference. Also, the proposed architecture allows propagating information about known labels directly to the intermediate layers to improve the final representation. Extensive evaluation of the proposed method confirms that our Plugin Networks outperform the state-of-the-art in a variety of tasks, including scene categorization, multilabel image annotation, and semantic segmentation.
Michal Koperski, Tomasz K. Konopczynski, Rafal Nowak, Piotr Semberecki, Tomasz Trzcinski
WACV5
2020 Adversarial autoencoders for compact representations of 3D point clouds
Maciej Zamorski, Maciej Zieba, Piotr Klukowski, Rafal Nowak, Karol Kurach, Wojciech Stokowiec, Tomasz Trzcinski
Comput. Vis. Image Underst.7
2019 Comixify: Transform Video Into Comics
abstract
In this paper, we propose a solution to transform a video into a comics. We approach this task using a neural style algorithm based on Generative Adversarial Networks (GANs). Several recent works in the field of Neural Style Transfer showed that producing an image in the style of another image is f easible. In this paper, we build up on these works and extend the existing set of style transfer use cases with a working application of video comixification. To that end, we train an end-to-end solution that transforms input video into a comics in two stages. In the first stage, we propose a state-of-the-art keyframes extraction algorithm that selects a subset of frames from the video to provide the most comprehensive video context and we filter those frames using image aesthetic estimation engine. In the second stage, the style of selected keyframes is transferred into a comics. To provide the most aesthetically compelling results, we selected the most state-of-the art style transfer solution and based on that implement our own ComixGAN framework. The final contribution of our work is a Web-based working application of video comixification available at http://comixify.ii.pw.edu.pl.
Maciej Pesko, Adam Svystun, Pawel Andruszkiewicz, Przemyslaw Rokita, Tomasz Trzcinski
Fundam. Informaticae5
2019 Classifying and Visualizing Emotions with Emotional DAN
abstract
Classification of human emotions remains an important and challenging task for many computer vision algorithms, especially in the era of humanoid robots which coexist with humans in their everyday life. Currently proposed methods for emotion recognition solve this task using multi-layered convoluti onal networks that do not explicitly infer any facial features in the classification phase. In this work, we postulate a fundamentally different approach to solve emotion recognition task that relies on incorporating facial landmarks as a part of the classification loss function. To that end, we extend a recently proposed Deep Alignment Network (DAN) with a term related to facial features. Thanks to this simple modification, our model called EmotionalDAN is able to outperform state-of-the-art emotion classification methods on two challenging benchmark dataset by up to 5%. Furthermore, we visualize image regions analyzed by the network when making a decision and the results indicate that our EmotionalDAN model is able to correctly identify facial landmarks responsible for expressing the emotions.
Ivona Tautkute, Tomasz Trzcinski
Fundam. Informaticae2
2018 Siamese Generative Adversarial Privatizer for Biometric Data
Witold Oleszkiewicz, Peter Kairouz, Karol J. Piczak, Ram Rajagopal, Tomasz Trzcinski
ACCV (5)5
2018 Estimating Achilles Tendon Healing Progress with Convolutional Neural Networks
Norbert Kapinski, Jakub Zielinski, Bartosz A. Borucki, Tomasz Trzcinski, Beata Ciszkowska-Lyson, Krzysztof Nowinski
MICCAI (2)4
2018 BinGAN: Learning Compact Binary Descriptors with a Regularized GAN
abstract
In this paper, we propose a novel regularization method for Generative Adversarial Networks that allows the model to learn discriminative yet compact binary representations of image patches (image descriptors). We exploit the dimensionality reduction that takes place in the intermediate layers of the discriminator network and train the binarized penultimate layer's low-dimensional representation to mimic the distribution of the higher-dimensional preceding layers. To achieve this, we introduce two loss terms that aim at: (i) reducing the correlation between the dimensions of the binarized penultimate layer's low-dimensional representation (i.e. maximizing joint entropy) and (ii) propagating the relations between the dimensions in the high-dimensional space to the low-dimensional space. We evaluate the resulting binary image descriptors on two challenging applications, image matching and retrieval, where they achieve state-of-the-art results.
Maciej Zieba, Piotr Semberecki, Tarek El-Gaaly, Tomasz Trzcinski
NeurIPS4
2017 What Looks Good with my Sofa: Ensemble Multimodal Search for Interior Design
abstract
In this paper, we propose a multi-modal search engine for interior design that combines visual and textual queries.The goal of our engine is to retrieve interior objects, e.g.furniture or wall clocks, that share visual and aesthetic similarities with the query.Our search engine allows the user to take a photo of a room and retrieve with a high recall a list of items identical or visually similar to those present in the photo.Additionally, it allows to return other items that aesthetically and stylistically fit well together.To achieve this goal, our system blends the results obtained using textual and visual modalities.Thanks to this blending strategy, we increase the average style similarity score of the retrieved items by 11%.Our work is implemented as a Web-based application and it is planned to be opened to the public.
Ivona Tautkute, Aleksandra Mozejko, Wojciech Stokowiec, Tomasz Trzcinski, Lukasz Brocki, Krzysztof Marasek
FedCSIS4
2017 Shallow Reading with Deep Learning: Predicting Popularity of Online Content Using only Its Title
Wojciech Stokowiec, Tomasz Trzcinski, Krzysztof Wolk, Krzysztof Marasek, Przemyslaw Rokita
ISMIS2
2017 Recurrent Neural Networks for Online Video Popularity Prediction
Tomasz Trzcinski, Pawel Andruszkiewicz, Tomasz Bochenski, Przemyslaw Rokita
ISMIS1
2017 Predicting Popularity of Online Videos Using Support Vector Regression
abstract
In this work, we propose a regression method to predict the popularity of an online video measured by its number of views. Our method uses Support Vector Regression with Gaussian radial basis functions. We show that predicting popularity patterns with this approach provides more precise and more stable prediction results, mainly thanks to the nonlinear character of the proposed method as well as its robustness. We prove the superiority of our method against the state of the art using datasets containing almost 24 000 videos from YouTube and Facebook. We also show that using visual features, such as the outputs of deep neural networks or scene dynamics' metrics, can be useful for popularity prediction before content publication. Furthermore, we show that popularity prediction accuracy can be improved by combining early distribution patterns with social and visual features and that social features represent a much stronger signal in terms of video popularity prediction than the visual ones.
Tomasz Trzcinski, Przemyslaw Rokita
IEEE Trans. Multim.1
2015 Learning Image Descriptors with Boosting
abstract
We propose a novel and general framework to learn compact but highly discriminative floating-point and binary local feature descriptors. By leveraging the boosting-trick we first show how to efficiently train a compact floating-point descriptor that is very robust to illumination and viewpoint changes. We then present the main contribution of this paper-a binary extension of the framework that demonstrates the real advantage of our approach and allows us to compress the descriptor even further. Each bit of the resulting binary descriptor, which we call BinBoost, is computed with a boosted binary hash function, and we show how to efficiently optimize the hash functions so that they are complementary, which is key to compactness and robustness. As we do not put any constraints on the weak learner configuration underlying each hash function, our general framework allows us to optimize the sampling patterns of recently proposed hand-crafted descriptors and significantly improve their performance. Moreover, our boosting scheme can easily adapt to new applications and generalize to other types of image data, such as faces, while providing state-of-the-art results at a fraction of the matching time and memory footprint.
Tomasz Trzcinski, C. Mario Christoudias, Vincent Lepetit
IEEE Trans. Pattern Anal. Mach. Intell.1
2014 Receptive Fields Selection for Binary Feature Description
abstract
Feature description for local image patch is widely used in computer vision. While the conventional way to design local descriptor is based on expert experience and knowledge, learning-based methods for designing local descriptor become more and more popular because of their good performance and data-driven property. This paper proposes a novel data-driven method for designing binary feature descriptor, which we call receptive fields descriptor (RFD). Technically, RFD is constructed by thresholding responses of a set of receptive fields, which are selected from a large number of candidates according to their distinctiveness and correlations in a greedy way. Using two different kinds of receptive fields (namely rectangular pooling area and Gaussian pooling area) for selection, we obtain two binary descriptors RFDR and RFDG .accordingly. Image matching experiments on the well-known patch data set and Oxford data set demonstrate that RFD significantly outperforms the state-of-the-art binary descriptors, and is comparable with the best float-valued descriptors at a fraction of processing time. Finally, experiments on object recognition tasks confirm that both RFDR and RFDG successfully bridge the performance gap between binary descriptors and their floating-point competitors.
Bin Fan 0001, Qingqun Kong, Tomasz Trzcinski, Zhiheng Wang 0001, Chunhong Pan, Pascal Fua
IEEE Trans. Image Process.3
2013 Boosting Binary Keypoint Descriptors
abstract
Binary key point descriptors provide an efficient alternative to their floating-point competitors as they enable faster processing while requiring less memory. In this paper, we propose a novel framework to learn an extremely compact binary descriptor we call Bin Boost that is very robust to illumination and viewpoint changes. Each bit of our descriptor is computed with a boosted binary hash function, and we show how to efficiently optimize the different hash functions so that they complement each other, which is key to compactness and robustness. The hash functions rely on weak learners that are applied directly to the image patches, which frees us from any intermediate representation and lets us automatically learn the image gradient pooling configuration of the final descriptor. Our resulting descriptor significantly outperforms the state-of-the-art binary descriptors and performs similarly to the best floating-point descriptors at a fraction of the matching time and memory footprint.
Tomasz Trzcinski, C. Mario Christoudias, Pascal Fua, Vincent Lepetit
CVPR1
2012 Efficient Discriminative Projections for Compact Binary Descriptors
Tomasz Trzcinski, Vincent Lepetit
ECCV (1)1
2012 Learning Image Descriptors with the Boosting-Trick
abstract
In this paper we apply boosting to learn complex non-linear local visual feature representations, drawing inspiration from its successful application to visual object detection. The main goal of local feature descriptors is to distinctively represent a salient image region while remaining invariant to viewpoint and illumination changes. This representation can be improved using machine learning, however, past approaches have been mostly limited to learning linear feature mappings in either the original input or a kernelized input feature space. While kernelized methods have proven somewhat effective for learning non-linear local feature descriptors, they rely heavily on the choice of an appropriate kernel function whose selection is often difficult and non-intuitive. We propose to use the boosting-trick to obtain a non-linear mapping of the input to a high-dimensional feature space. The non-linear feature mapping obtained with the boosting-trick is highly intuitive. We employ gradient-based weak learners resulting in a learned descriptor that closely resembles the well-known SIFT. As demonstrated in our experiments, the resulting descriptor can be learned directly from intensity patches achieving state-of-the-art performance.
Tomasz Trzcinski, C. Mario Christoudias, Vincent Lepetit, Pascal Fua
NIPS1
2012 BRIEF: Computing a Local Binary Descriptor Very Fast
abstract
Binary descriptors are becoming increasingly popular as a means to compare feature points very fast while requiring comparatively small amounts of memory. The typical approach to creating them is to first compute floating-point ones, using an algorithm such as SIFT, and then to binarize them. In this paper, we show that we can directly compute a binary descriptor, which we call BRIEF, on the basis of simple intensity difference tests. As a result, BRIEF is very fast both to build and to match. We compare it against SURF and SIFT on standard benchmarks and show that it yields comparable recognition accuracy, while running in an almost vanishing fraction of the time required by either.
Michael Calonder, Vincent Lepetit, Mustafa Özuysal, Tomasz Trzcinski, Christoph Strecha, Pascal Fua
IEEE Trans. Pattern Anal. Mach. Intell.4
2012 Thick boundaries in binary space and their influence on nearest-neighbor search
Tomasz Trzcinski, Vincent Lepetit, Pascal Fua
Pattern Recognit. Lett.1