EDBT 2026 Demo / reviewers in the wild / expert
Mateusz Kozinski
dblp:122/8721
· DBLP profile ↗
23ranked-venue papers
5as first author
14since 2021 · last 2024
0000-0002-3187-518XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | MULDE: Multiscale Log-Density Estimation via Denoising Score Matching for Video Anomaly DetectionabstractWe propose a novel approach to video anomaly detection: we treat feature vectors extracted from videos as re-alizations of a random variable with a fixed distribution and model this distribution with a neural network. This lets us estimate the likelihood of test videos and detect video anomalies by thresholding the likelihood estimates. We train our video anomaly detector using a modification of de-noising score matching, a method that injects training data with noise to facilitate modeling its distribution. To elim-inate hyperparameter selection, we model the distribution of noisy video features across a range of noise levels and introduce a regularizer that tends to align the models for different levels of noise. At test time, we combine anomaly indications at multiple noise scales with a Gaussian mix-ture model. Running our video anomaly detector induces minimal delays as inference requires merely extracting the features and forward-propagating them through a shallow neural network and a Gaussian mixture model. Our ex-periments on five popular video anomaly detection bench-marks demonstrate state-of-the-art performance, both in the object-centric and in the frame-centric setup. Jakub Micorek, Horst Possegger, Dominik Narnhofer, Horst Bischof, Mateusz Kozinski |
CVPR | 5 |
| 2024 | Meta-prompting for Automating Zero-Shot Visual Recognition with LLMs
Muhammad Jehanzeb Mirza, Leonid Karlinsky, Wei Lin 0019, Sivan Doveh, Jakub Micorek, Mateusz Kozinski, Hilde Kuehne, Horst Possegger |
ECCV (2) | 6 |
| 2024 | Error Management for Augmented Reality Assembly InstructionsabstractAugmented reality (AR) lends itself to presenting visual instructions on how to assemble or disassemble an object. Splitting the assembly procedure into shorter steps and presenting the corresponding instructions in AR supports their comprehension. However, one can still misinterpret instructions and make errors while manipulating the object. While previous work supports detecting the occurrence of errors, we investigate handling such errors. This requires knowledge of the error at runtime of the application. Starting from a categorization of the errors, we investigate how to automatically derive common error states to generate training data. We introduce an extension to a state-of-the-art deep-learning-based object detector for supporting the detection of assembly states at real-time update rates, based on contrastive learning. We evaluated the proposed detector, showing that it outperforms the state-of-the-art, and we demonstrate our work with an AR application that alerts the user if errors occur and provides visual help to correct the error. Ana Stanescu 0003, Peter Mohr, Franz Thaler, Mateusz Kozinski, Lucchas Ribeiro Skreinig, Dieter Schmalstieg, Denis Kalkofen |
ISMAR | 4 |
| 2023 | Video Test-Time Adaptation for Action RecognitionabstractAlthough action recognition systems can achieve top performance when evaluated on in-distribution test points, they are vulnerable to unanticipated distribution shifts in test data. However, test-time adaptation of video action recognition models against common distribution shifts has so far not been demonstrated. We propose to address this problem with an approach tailored to spatio-temporal models that is capable of adaptation on a single video sample at a step. It consists in a feature distribution alignment technique that aligns online estimates of test set statistics towards the training statistics. We further enforce prediction consistency over temporally augmented views of the same test video sample. Evaluations on three benchmark action recognition datasets show that our proposed technique is architecture-agnostic and able to significantly boost the performance on both, the state of the art convolutional architecture TANet and the Video Swin Transformer. Our proposed method demonstrates a substantial performance gain over existing test-time adaptation approaches in both evaluations of a single distribution shift and the challenging case of random distribution shifts. Code will be available at https://github.com/wlin-at/ViTTA. Wei Lin 0019, Muhammad Jehanzeb Mirza, Mateusz Kozinski, Horst Possegger, Hilde Kuehne, Horst Bischof |
CVPR | 3 |
| 2023 | ActMAD: Activation Matching to Align Distributions for Test-Time-TrainingabstractTest-Time-Training (TTT) is an approach to cope with out-of-distribution (OOD) data by adapting a trained model to distribution shifts occurring at test-time. We propose to perform this adaptation via Activation Matching (ActMAD): We analyze activations of the model and align activation statistics of the OOD test data to those of the training data. In contrast to existing methods, which model the distribution of entire channels in the ultimate layer of the feature extractor, we model the distribution of each feature in multiple layers across the network. This results in a more fine-grained supervision and makes ActMAD attain state of the art performance on CIFAR-100C and Imagenet-C. ActMAD is also architecture-and task-agnostic, which lets us go beyond image classification, and score 15.4% improvement over previous approaches when evaluating a KITTI-trained object detector on KITTI-Fog. Our experiments highlight that ActMAD can be applied to online adaptation in realistic scenarios, requiring little data to attain its full performance. Muhammad Jehanzeb Mirza, Pol Jané-Soneira, Wei Lin 0019, Mateusz Kozinski, Horst Possegger, Horst Bischof |
CVPR | 4 |
| 2023 | MAtch, eXpand and Improve: Unsupervised Finetuning for Zero-Shot Action Recognition with Language KnowledgeabstractLarge scale Vision Language (VL) models have shown tremendous success in aligning representations between visual and text modalities. This enables remarkable progress in zero-shot recognition, image generation & editing, and many other exciting tasks. However, VL models tend to over-represent objects while paying much less attention to verbs, and require additional tuning on video data for best zero-shot action recognition performance. While previous work relied on large-scale, fully-annotated data, in this work we propose an unsupervised approach. We adapt a VL model for zero-shot and few-shot action recognition using a collection of unlabeled videos and an unpaired action dictionary. Based on that, we leverage Large Language Models and VL models to build a text bag for each unlabeled video via matching, text expansion and captioning. We use those bags in a Multiple Instance Learning setup to adapt an image-text backbone to video data. Although finetuned on unlabeled video data, our resulting models demonstrate high transferability to numerous unseen zero-shot downstream tasks, improving the base VL model performance by up to 14%, and even comparing favorably to fully-supervised baselines in both zero-shot and few-shot video recognition transfer. The code is released at https://github.com/wlin-at/MAXI. Wei Lin 0019, Leonid Karlinsky, Nina Shvetsova, Horst Possegger, Mateusz Kozinski, Rameswar Panda, Rogério Feris, Hilde Kuehne, Horst Bischof |
ICCV | 5 |
| 2023 | MATE: Masked Autoencoders are Online 3D Test-Time LearnersabstractOur MATE is the first Test-Time-Training (TTT) method designed for 3D data, which makes deep networks trained for point cloud classification robust to distribution shifts occurring in test data. Like existing TTT methods from the 2D image domain, MATE also leverages test data for adaptation. Its test-time objective is that of a Masked Autoencoder: a large portion of each test point cloud is removed before it is fed to the network, tasked with reconstructing the full point cloud. Once the network is updated, it is used to classify the point cloud. We test MATE on several 3D object classification datasets and show that it significantly improves robustness of deep networks to several types of corruptions commonly occurring in 3D point clouds. We show that MATE is very efficient in terms of the fraction of points it needs for the adaptation. It can effectively adapt given as few as 5% of tokens of each test sample, making it extremely lightweight. Our experiments show that MATE also achieves competitive performance by adapting sparsely on the test data, which further reduces its computational overhead, making it ideal for real-time applications. Muhammad Jehanzeb Mirza, Inkyu Shin, Wei Lin 0019, Andreas Schriebl, Kunyang Sun, Jaesung Choe, Mateusz Kozinski, Horst Possegger, In-So Kweon, Kuk-Jin Yoon, Horst Bischof |
ICCV | 7 |
| 2023 | State-Aware Configuration Detection for Augmented Reality Step-by-Step TutorialsabstractPresenting tutorials in augmented reality is a compelling application area, but previous attempts have been limited to objects with only a small numbers of parts. Scaling augmented reality tutorials to complex assemblies of a large number of parts is difficult, because it requires automatically discriminating many similar-looking object configurations, which poses a challenge for current object detection techniques. In this paper, we seek to lift this limitation. Our approach is inspired by the observation that, even though the number of assembly steps may be large, their order is typically highly restricted: Some actions can only be performed after others. To leverage this observation, we enhance a state-of-the-art object detector to predict the current assembly state by conditioning on the previous one, and to learn the constraints on consecutive states. This learned ‘consecutive state prior’ helps the detector disambiguate configurations that are otherwise too similar in terms of visual appearance to be reliably discriminated. Via the state prior, the detector is also able to improve the estimated probabilities that a state detection is correct. We experimentally demonstrate that our technique enhances the detection accuracy for assembly sequences with a large number of steps and on a variety of use cases, including furniture, Lego and origami. Additionally, we demonstrate the use of our algorithm in an interactive augmented reality application. Ana Stanescu 0003, Peter Mohr, Mateusz Kozinski, Shohei Mori, Dieter Schmalstieg, Denis Kalkofen |
ISMAR | 3 |
| 2023 | Sit Back and Relax: Learning to Drive Incrementally in All Weather ConditionsabstractIn autonomous driving scenarios, current object detection models show strong performance when tested in clear weather. However, their performance deteriorates significantly when tested in degrading weather conditions. In addition, even when adapted to perform robustly in a sequence of different weather conditions, they are often unable to perform well in all of them and suffer from catastrophic forgetting. To efficiently mitigate forgetting, we propose Domain-Incremental Learning through Activation Matching (DILAM), which employs unsupervised feature alignment to adapt only the affine parameters of a clear weather pre-trained network to different weather conditions. We propose to store these affine parameters as a memory bank for each weather condition and plug-in their weather-specific parameters during driving (i.e. test time) when the respective weather conditions are encountered. Our memory bank is extremely lightweight, since affine parameters account for less than 2% of a typical object detector. Furthermore, contrary to previous domain-incremental learning approaches, we do not require the weather label when testing and propose to automatically infer the weather condition by a majority voting linear classifier. Stefan Leitner, Muhammad Jehanzeb Mirza, Wei Lin 0019, Jakub Micorek, Marc Masana, Mateusz Kozinski, Horst Possegger, Horst Bischof |
IV | 6 |
| 2023 | LaFTer: Label-Free Tuning of Zero-shot Classifier using Language and Unlabeled Image CollectionsabstractRecently, large-scale pre-trained Vision and Language (VL) models have set a new state-of-the-art (SOTA) in zero-shot visual classification enabling open-vocabulary recognition of potentially unlimited set of categories defined as simple language prompts. However, despite these great advances, the performance of these zero-shot classifiers still falls short of the results of dedicated (closed category set) classifiers trained with supervised fine-tuning. In this paper we show, for the first time, how to reduce this gap without any labels and without any paired VL data, using an unlabeled image collection and a set of texts auto-generated using a Large Language Model (LLM) describing the categories of interest and effectively substituting labeled visual instances of those categories. Using our label-free approach, we are able to attain significant performance improvements over the zero-shot performance of the base VL model and other contemporary methods and baselines on a wide variety of datasets, demonstrating absolute improvement of up to $11.7\%$ ($3.8\%$ on average) in the label-free setting. Moreover, despite our approach being label-free, we observe $1.3\%$ average gains over leading few-shot prompting baselines that do use 5-shot supervision. Muhammad Jehanzeb Mirza, Leonid Karlinsky, Wei Lin 0019, Horst Possegger, Mateusz Kozinski, Rogério Feris, Horst Bischof |
NeurIPS | 5 |
| 2023 | Persistent Homology With Improved Locality Information for More Effective DelineationabstractPersistent Homology (PH) has been successfully used to train networks to detect curvilinear structures and to improve the topological quality of their results. However, existing methods are very global and ignore the location of topological features. In this paper, we remedy this by introducing a new filtration function that fuses two earlier approaches: thresholding-based filtration, previously used to train deep networks to segment medical images, and filtration with height functions, typically used to compare 2D and 3D shapes. We experimentally demonstrate that deep networks trained using our PH-based loss function yield reconstructions of road networks and neuronal processes that reflect ground-truth connectivity better than networks trained with existing loss functions based on PH. Doruk Öner, Adélie Garin, Mateusz Kozinski, Kathryn Hess, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Enforcing Connectivity of 3D Linear Structures Using Their 2D Projections
Doruk Öner, Hussein Osman, Mateusz Kozinski, Pascal Fua |
MICCAI (5) | 3 |
| 2022 | Promoting Connectivity of Network-Like Structures by Enforcing Region SeparationabstractWe propose a novel, connectivity-oriented loss function for training deep convolutional networks to reconstruct network-like structures, like roads and irrigation canals, from aerial images. The main idea behind our loss is to express the connectivity of roads, or canals, in terms of disconnections that they create between background regions of the image. In simple terms, a gap in the predicted road causes two background regions, that lie on the opposite sides of a ground truth road, to touch in prediction. Our loss function is designed to prevent such unwanted connections between background regions, and therefore close the gaps in predicted roads. It also prevents predicting false positive roads and canals by penalizing unwarranted disconnections of background regions. In order to capture even short, dead-ending road segments, we evaluate the loss in small image crops. We show, in experiments on two standard road benchmarks and a new data set of irrigation canals, that convnets trained with our loss function recover road connectivity so well that it suffices to skeletonize their output to produce state of the art maps. A distinct advantage of our approach is that the loss can be plugged in to any existing training setup without further modifications. Doruk Öner, Mateusz Kozinski, Leonardo Citraro, Nathan C. Dadap, Alexandra Georges Konings, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Adjusting the Ground Truth Annotations for Connectivity-Based Learning to DelineateabstractDeep learning-based approaches to delineating 3D structure depend on accurate annotations to train the networks. Yet in practice, people, no matter how conscientious, have trouble precisely delineating in 3D and on a large scale, in part because the data is often hard to interpret visually and in part because the 3D interfaces are awkward to use. In this paper, we introduce a method that explicitly accounts for annotation inaccuracies. To this end, we treat the annotations as active contour models that can deform themselves while preserving their topology. This enables us to jointly train the network and correct potential errors in the original annotations. The result is an approach that boosts performance of deep networks trained with potentially inaccurate annotations. Doruk Öner, Mateusz Kozinski, Leonardo Citraro, Pascal Fua |
IEEE Trans. Medical Imaging | 2 |
| 2020 | Towards Reliable Evaluation of Algorithms for Road Network Reconstruction from Aerial Images
Leonardo Citraro, Mateusz Kozinski, Pascal Fua |
ECCV (28) | 2 |
| 2020 | TopoAL: An Adversarial Learning Approach for Topology-Aware Road Segmentation
Subeesh Vasu, Mateusz Kozinski, Leonardo Citraro, Pascal Fua |
ECCV (27) | 2 |
| 2020 | Tracing in 2D to reduce the annotation effort for 3D deep delineation of linear structures
Mateusz Kozinski, Agata Mosinska, Mathieu Salzmann, Pascal Fua |
Medical Image Anal. | 1 |
| 2020 | Joint Segmentation and Path Classification of Curvilinear StructuresabstractDetection of curvilinear structures in images has long been of interest. One of the most challenging aspects of this problem is inferring the graph representation of the curvilinear network. Most existing delineation approaches first perform binary segmentation of the image and then refine it using either a set of hand-designed heuristics or a separate classifier that assigns likelihood to paths extracted from the pixel-wise prediction. In our work, we bridge the gap between segmentation and path classification by training a deep network that performs those two tasks simultaneously. We show that this approach is beneficial because it enforces consistency across the whole processing pipeline. We apply our approach on roads and neurons datasets. Agata Mosinska, Mateusz Kozinski, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Beyond the Pixel-Wise Loss for Topology-Aware DelineationabstractDelineation of curvilinear structures is an important problem in Computer Vision with multiple practical applications. With the advent of Deep Learning, many current approaches on automatic delineation have focused on finding more powerful deep architectures, but have continued using the habitual pixel-wise losses such as binary cross-entropy. In this paper we claim that pixel-wise losses alone are unsuitable for this problem because of their inability to reflect the topological impact of mistakes in the final prediction. We propose a new loss term that is aware of the higher-order topological features of linear structures. We also exploit a refinement pipeline that iteratively applies the same model over the previous delineation to refine the predictions at each step, while keeping the number of parameters and the complexity of the model constant. When combined with the standard pixel-wise loss, both our new loss term and an iterative refinement boost the quality of the predicted delineations, in some cases almost doubling the accuracy as compared to the same classifier trained with the binary cross-entropy alone. We show that our approach outperforms state-of-the-art methods on a wide range of data, from microscopy to aerial images. Agata Mosinska, Pablo Márquez-Neila, Mateusz Kozinski, Pascal Fua |
CVPR | 3 |
| 2018 | Learning to Segment 3D Linear Structures Using Only 2D Annotations
Mateusz Kozinski, Agata Mosinska, Mathieu Salzmann, Pascal Fua |
MICCAI (2) | 1 |
| 2015 | A MRF shape prior for facade parsing with occlusionsabstractWe present a new shape prior formalism for the segmentation of rectified facade images. It combines the simplicity of split grammars with unprecedented expressive power: the capability of encoding simultaneous alignment in two dimensions, facade occlusions and irregular boundaries between facade elements. We formulate the task of finding the most likely image segmentation conforming to a prior of the proposed form as a MAP-MRF problem over a 4-connected pixel grid, and propose an efficient optimization algorithm for solving it. Our method simultaneously segments the visible and occluding objects, and recovers the structure of the occluded facade. We demonstrate state-of-the-art results on a number of facade segmentation datasets. Mateusz Kozinski, Raghudeep Gadde, Sergey Zagoruyko, Guillaume Obozinski, Renaud Marlet |
CVPR | 1 |
| 2014 | Beyond Procedural Facade Parsing: Bidirectional Alignment via Linear Programming
Mateusz Kozinski, Guillaume Obozinski, Renaud Marlet |
ACCV (4) | 1 |
| 2014 | Image parsing with graph grammars and Markov Random Fields applied to facade analysisabstractExisting approaches to parsing images of objects featuring complex, non-hierarchical structure rely on exploration of a large search space combining the structure of the object and positions of its parts. The latter task requires randomized or greedy algorithms that do not produce repeatable results or strongly depend on the initial solution. To address the problem we propose to model and optimize the structure of the object and position of its parts separately. We encode the possible object structures in a graph grammar. Then, for a given structure, the positions of the parts are inferred using standard MAP-MRF techniques. This way we limit the application of the less reliable greedy or randomized optimization algorithm to structure inference. We apply our method to parsing images of building facades. The results of our experiments compare favorably to the state of the art. Mateusz Kozinski, Renaud Marlet |
WACV | 1 |