EDBT 2026 Demo / reviewers in the wild / expert
Sinisa Segvic
dblp:42/3275
· DBLP profile ↗
36ranked-venue papers
7as first author
19since 2021 · last 2025
0000-0001-7378-0536ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 5 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 5 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Systems, architecture and hardware · 2Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Sequential Keypoint Density Estimator: an Overlooked Baseline of Skeleton-Based Video Anomaly Detection
Anja Delic, Matej Grcic, Sinisa Segvic |
ICCV | 3 |
| 2025 | Seal Your Backdoor with Variational DefenseabstractWe propose VIBE, a model-agnostic framework that trains classifiers resilient to backdoor attacks. The key concept behind our approach is to treat malicious inputs and corrupted labels from the training dataset as observed random variables, while the actual clean labels are latent. VIBE then recovers the corresponding latent clean label posterior through variational inference. The resulting training procedure follows the expectation-maximization (EM) algorithm. The E-step infers the clean pseudolabels by solving an entropy-regularized optimal transport problem, while the M-step updates the classifier parameters via gradient descent. Being modular, VIBE can seamlessly integrate with recent advancements in self-supervised representation learning, which enhance its ability to resist backdoor attacks. We experimentally validate the method effectiveness against contemporary backdoor attacks on standard datasets, a large-scale setup with 1$k$ classes, and a dataset poisoned with multiple attacks. VIBE consistently outperforms previous defenses across all tested scenarios. Ivan Sabolic, Matej Grcic, Sinisa Segvic |
ICCV | 3 |
| 2024 | Quantile-Based Maximum Likelihood Training for Outlier DetectionabstractDiscriminative learning effectively predicts true object class for image classification. However, it often results in false positives for outliers, posing critical concerns in applications like autonomous driving and video surveillance systems. Previous attempts to address this challenge involved training image classifiers through contrastive learning using actual outlier data or synthesizing outliers for self-supervised learning. Furthermore, unsupervised generative modeling of inliers in pixel space has shown limited success for outlier detection. In this work, we introduce a quantile-based maximum likelihood objective for learning the inlier distribution to improve the outlier separation during inference. Our approach fits a normalizing flow to pre-trained discriminative features and detects the outliers according to the evaluated log-likelihood. The experimental evaluation demonstrates the effectiveness of our method as it surpasses the performance of the state-of-the-art unsupervised methods for outlier detection. The results are also competitive compared with a recent self-supervised approach for outlier detection. Our work allows to reduce dependency on well-sampled negative training data, which is especially important for domains like medical diagnostics or remote sensing. Masoud Taghikhah, Nishant Kumar 0005, Sinisa Segvic, Abouzar Eslami, Stefan Gumhold |
AAAI | 3 |
| 2024 | Outlier detection by ensembling uncertainty with negative objectness
Anja Delic, Matej Grcic, Sinisa Segvic |
BMVC | 3 |
| 2024 | Backdoor Defense through Self-Supervised and Generative Learning
Ivan Sabolic, Ivan Grubisic 0001, Sinisa Segvic |
BMVC | 3 |
| 2024 | MC-PanDA: Mask Confidence for Panoptic Domain Adaptation
Ivan Martinovic, Josip Saric, Sinisa Segvic |
ECCV (72) | 3 |
| 2024 | Identifying Label Errors in Object Detection Datasets by Loss InspectionabstractLabeling datasets for supervised object detection is a dull and time-consuming task. Errors can be easily introduced during annotation and overlooked during review, yielding inaccurate benchmarks and performance degradation of deep neural networks trained on noisy labels. In this work, we introduce a benchmark for label error detection methods on object detection datasets as well as a theoretically underpinned label error detection method and a number of baselines. We simulate four different types of randomly introduced label errors on train and test sets of well-labeled object detection datasets. For our label error detection method we assume a two-stage object detector to be given and consider the sum of both stages’ classification and regression losses. The losses are computed with respect to the predictions and the noisy labels including simulated label errors, aiming at detecting the latter. We compare our method to four baselines: a naive one without deep learning, the object detector’s score, the entropy of the classification softmax distribution and a probability margin based method from related work. We outperform all baselines and demonstrate that among the considered methods, ours is the only one that detects label errors of all four types efficiently, which we also derive theoretically. Furthermore, we detect real label errors a) on commonly used test datasets in object detection and b) on a proprietary dataset. In both cases we achieve low false positives rates, i.e., we detect label errors with a precision for a) of up to 71.5% and for b) with 97%. Marius Schubert, Tobias Riedlinger, Karsten Kahl, Daniel Kröll, Sebastian Schoenen, Sinisa Segvic, Matthias Rottmann |
WACV | 6 |
| 2024 | Weakly Supervised Training of Universal Visual Concepts for Multi-domain Semantic Segmentation
Petra Bevandic, Marin Orsic, Josip Saric, Ivan Grubisic 0001, Sinisa Segvic |
Int. J. Comput. Vis. | 5 |
| 2024 | Hybrid Open-Set Segmentation With Synthetic Negative DataabstractOpen-set segmentation can be conceived by complementing closed-set classification with anomaly detection. Many of the existing dense anomaly detectors operate through generative modelling of regular data or by discriminating with respect to negative data. These two approaches optimize different objectives and therefore exhibit different failure modes. Consequently, we propose a novel anomaly score that fuses generative and discriminative cues. Our score can be implemented by upgrading any closed-set segmentation model with dense estimates of dataset posterior and unnormalized data likelihood. The resulting dense hybrid open-set models require negative training images that can be sampled from an auxiliary negative dataset, from a jointly trained generative model, or from a mixture of both sources. We evaluate our contributions on benchmarks for dense anomaly detection and open-set segmentation. The experiments reveal strong open-set performance in spite of negligible computational overhead. Matej Grcic, Sinisa Segvic |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Dynamic Loss Balancing and Sequential Enhancement for Road-Safety Assessment and Traffic Scene ClassificationabstractRoad-safety inspection is an indispensable instrument for reducing road-accident fatalities related to road infrastructure. Recent work formalizes the assessment procedure in terms of carefully selected risk factors that are also known as road-safety attributes. In current practice, these attributes are manually annotated in geo-referenced monocular video for each road segment. We propose to reduce dependency on tedious human labor by automating attribute collection through a two-stage deep learning approach. The first stage recognizes more than forty road-safety attributes by observing a local spatio-temporal context. Our design leverages an efficient convolutional pipeline, which benefits from pre-training on semantic segmentation of street scenes. The second stage enhances predictions through sequential integration across a larger temporal window. Our design leverages per-attribute instances of a lightweight recurrent architecture. Both stages alleviate extreme class imbalance by incorporating a multi-task variant of recall-based dynamic loss weighting. We perform experiments on the novel iRAP-BH dataset, which involves fully labeled geo-referenced video along 2,300 km of public roads in Bosnia and Herzegovina. Moreover, we evaluate our approach against the related work on three road-scene classification datasets from the literature: Honda Scenes, FM3m, and BDD100k. Experimental evaluation confirms the value of our contributions on all three datasets. Marin Kacan, Marko Sevrovic, Sinisa Segvic |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | Normalizing Flow based Feature Synthesis for Outlier-Aware Object DetectionabstractReal-world deployment of reliable object detectors is crucial for applications such as autonomous driving. However, general-purpose object detectors like Faster R-CNN are prone to providing overconfident predictions for outlier objects. Recent outlier-aware object detection approaches estimate the density of instance-wide features with class-conditional Gaussians and train on synthesized outlier features from their low-likelihood regions. However, this strategy does not guarantee that the synthesized outlier features will have a low likelihood according to the other class-conditional Gaussians. We propose a novel outlier-aware object detection framework that distinguishes outliers from inlier objects by learning the joint data distribution of all inlier classes with an invertible normalizing flow. The appropriate sampling of the flow model ensures that the synthesized outliers have a lower likelihood than inliers of all object classes, thereby modeling a better decision boundary between inlier and outlier objects. Our approach significantly outperforms the state-of-the-art for outlier-aware object detection on both image and video datasets. Nishant Kumar 0005, Sinisa Segvic, Abouzar Eslami, Stefan Gumhold |
CVPR | 2 |
| 2023 | Dense Semantic Forecasting in Video by Joint Regression of Features and Feature MotionabstractDense semantic forecasting anticipates future events in the video by inferring pixel-level semantics of an unobserved future image. We present a novel approach that is applicable to various single-frame architectures and tasks. Our approach consists of two modules. The feature-to-motion (F2M) module forecasts a dense deformation field that warps past features into their future positions. The feature-to-feature (F2F) module regresses the future features directly and is, therefore, able to account for emergent scenery. The compound F2MF model decouples the effects of motion from the effects of novelty in a task-agnostic manner. We aim to apply F2MF forecasting to the most subsampled and the most abstract representation of the desired single-frame model. Our design takes advantage of deformable convolutions and spatial correlation coefficients across neighboring time instants. We perform experiments on three dense prediction tasks: semantic segmentation, instance-level segmentation, and panoptic segmentation. The results reveal state-of-the-art forecasting accuracy across three dense prediction tasks. Josip Saric, Sacha Vrazic, Sinisa Segvic |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Automatic universal taxonomies for multi-domain semantic segmentation
Petra Bevandic, Sinisa Segvic |
BMVC | 2 |
| 2022 | DenseHybrid: Hybrid Anomaly Detection for Dense Open-Set Recognition
Matej Grcic, Petra Bevandic, Sinisa Segvic |
ECCV (25) | 3 |
| 2022 | Multi-domain semantic segmentation with overlapping labels *abstractDeep supervised models have an unprecedented capacity to absorb large quantities of training data. Hence, training on many datasets becomes a method of choice towards graceful degradation in unusual scenes. Unfortunately, different datasets often use incompatible labels. For instance, the Cityscapes road class subsumes all driving surfaces, while Vistas defines separate classes for road markings, manholes etc. We address this challenge by proposing a principled method for seamless learning on datasets with overlapping classes based on partial labels and probabilistic loss. Our method achieves competitive within-dataset and cross-dataset generalization, as well as ability to learn visual concepts which are not separately labeled in any of the training datasets. Experiments reveal competitive or state-of-the-art performance on two multi-domain dataset collections and on the WildDash 2 benchmark. Petra Bevandic, Marin Orsic, Ivan Grubisic 0001, Josip Saric, Sinisa Segvic |
WACV | 5 |
| 2022 | Dense open-set recognition based on training with noisy negative images
Petra Bevandic, Ivan Kreso, Marin Orsic, Sinisa Segvic |
Image Vis. Comput. | 4 |
| 2021 | Densely connected normalizing flowsabstractNormalizing flows are bijective mappings between inputs and latent representations with a fully factorized distribution. They are very attractive due to exact likelihood evaluation and efficient sampling. However, their effective capacity is often insufficient since the bijectivity constraint limits the model width. We address this issue by incrementally padding intermediate representations with noise. We precondition the noise in accordance with previous invertible units, which we describe as cross-unit coupling. Our invertible glow-like modules increase the model expressivity by fusing a densely connected block with Nyström self-attention. We refer to our architecture as DenseFlow since both cross-unit and intra-module couplings rely on dense connectivity. Experiments show significant improvements due to the proposed contributions and reveal state-of-the-art density estimation under moderate computing budgets. Matej Grcic, Ivan Grubisic 0001, Sinisa Segvic |
NeurIPS | 3 |
| 2021 | Efficient semantic segmentation with pyramidal fusion
Marin Orsic, Sinisa Segvic |
Pattern Recognit. | 2 |
| 2021 | Efficient Ladder-Style DenseNets for Semantic Segmentation of Large ImagesabstractRecent progress of deep image classification models provides great potential for improving related computer vision tasks. However, the transition to semantic segmentation is hampered by strict memory limitations of contemporary GPUs. The extent of feature map caching required by convolutional backprop poses significant challenges even for moderately sized Pascal images, while requiring careful architectural considerations when input resolution is in the megapixel range. To address these concerns, we propose a novel ladder-style DenseNet-based architecture which features high modelling power, efficient upsampling, and inherent spatial efficiency which we unlock with checkpointing. The resulting models deliver high performance and allow training at megapixel resolution on commodity hardware. The presented experimental results outperform the state-of-the-art in terms of prediction accuracy and execution speed on Cityscapes, VOC 2012, CamVid and ROB 2018 datasets. Source code at https://github.com/ivankreso/LDN. Ivan Kreso, Josip Krapac, Sinisa Segvic |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2020 | Warp to the Future: Joint Forecasting of Features and Feature MotionabstractWe address anticipation of scene development by forecasting semantic segmentation of future frames. Several previous works approach this problem by F2F (feature-to-feature) forecasting where future features are regressed from observed features. Different from previous work, we consider a novel F2M (feature-to-motion) formulation, which performs the forecast by warping observed features according to regressed feature flow. This formulation models a causal relationship between the past and the future, and regularizes inference by reducing dimensionality of the forecasting target. However, emergence of future scenery which was not visible in observed frames can not be explained by warping. We propose to address this issue by complementing F2M forecasting with the classic F2F approach. We realize this idea as a multi-head F2MF model built atop shared features. Experiments show that the F2M head prevails in static parts of the scene while the F2F head kicks-in to fill-in the novel regions. The proposed F2MF model operates in synergy with correlation features and outperforms all previous approaches both in short-term and mid-term forecast on the Cityscapes dataset. Josip Saric, Marin Orsic, Tonci Antunovic, Sacha Vrazic, Sinisa Segvic |
CVPR | 5 |
| 2020 | Traffic Scene Classification on a Representation BudgetabstractVisual cues can be used alongside GPS positioning and digital maps to improve understanding of vehicle environment in fleet management systems. Such systems are limited both in terms of bandwidth and storage space, so minimizing the size of transmitted and stored visual data is a priority. In this paper, we present efficient strategies for computing very short image representations suitable for classifying various types of traffic scenes in fleet management systems. We anticipate that the set of interesting classes will change over time, so we consider image representations that can be trained without knowing the labels of the target dataset. We empirically evaluate and compare the presented methods on a contributed dataset of 11447 labeled traffic scenes. Our results indicate that excellent classification results can be achieved with very short image representations, and that fine-tuning on the target dataset image data is not mandatory. Image descriptors can be as short as 128 components while still offering good performance, even in presence of adverse weather or illumination conditions. Ivan Sikiric, Karla Brkic, Petra Bevandic, Ivan Kreso, Josip Krapac, Sinisa Segvic |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2019 | In Defense of Pre-Trained ImageNet Architectures for Real-Time Semantic Segmentation of Road-Driving ImagesabstractRecent success of semantic segmentation approaches on demanding road driving datasets has spurred interest in many related application fields. Many of these applications involve real-time prediction on mobile platforms such as cars, drones and various kinds of robots. Real-time setup is challenging due to extraordinary computational complexity involved. Many previous works address the challenge with custom lightweight architectures which decrease computational complexity by reducing depth, width and layer capacity with respect to general purpose architectures. We propose an alternative approach which achieves a significantly better performance across a wide range of computing budgets. First, we rely on a light-weight general purpose architecture as the main recognition engine. Then, we leverage light-weight upsampling with lateral connections as the most cost-effective solution to restore the prediction resolution. Finally, we propose to enlarge the receptive field by fusing shared features at multiple resolutions in a novel fashion. Experiments on several road driving datasets show a substantial advantage of the proposed approach, either with ImageNet pre-trained parameters or when we learn from scratch. Our Cityscapes test submission entitled SwiftNetRN-18 delivers 75.5% MIoU and achieves 39.9 Hz on 1024×2048 images on GTX1080Ti. Marin Orsic, Ivan Kreso, Petra Bevandic, Sinisa Segvic |
CVPR | 4 |
| 2019 | Pedestrian Tracking by Probabilistic Data Association and Correspondence Embeddings
Borna Bicanic, Marin Orsic, Ivan Markovic, Sinisa Segvic, Ivan Petrovic |
FUSION | 4 |
| 2018 | Sparse weakly supervised models for object localization in road environment
Valentina Zadrija, Josip Krapac, Sinisa Segvic, Jakob Verbeek |
Comput. Vis. Image Underst. | 3 |
| 2014 | Image representations on a budget: Traffic scene classification in a restricted bandwidth scenarioabstractModern fleet management systems typically monitor the status of hundreds of vehicles by relying on GPS and other simple sensors. Such systems experience significant problems in cases of GPS glitches as well as in areas without GPS coverage. Additionally, when the tracked vehicle is stationary, they cannot discriminate between traffic jams, service stations, parking lots, serious accidents and other interesting scenarios. We propose to alleviate these problems by augmenting the GPS information with a short descriptor of an image captured by an on-board camera. The descriptor allows the server to recognize various scene types by image classification and to subsequently implement suitable business policies. Due to restricted bandwidth we focus on finding a compact image representation that would still allow reliable classification. We therefore consider several state-of-the-art descriptors under tight representation budgets of 512, 256, 128 and 64 components, and evaluate classification performance on a novel image dataset specifically crafted for fleet management applications. Experimental results indicate fair performance even with very short descriptor sizes and encourage further research in the field. Ivan Sikiric, Karla Brkic, Josip Krapac, Sinisa Segvic |
Intelligent Vehicles Symposium | 4 |
| 2014 | Exploiting temporal and spatial constraints in traffic sign detection from a moving vehicle
Sinisa Segvic, Karla Brkic, Zoran Kalafatic, Axel Pinz |
Mach. Vis. Appl. | 1 |
| 2011 | Experimental Evaluation of Autonomous Driving Based on Visual Memory and Image-Based Visual ServoingabstractIn this paper, the performance of a topological–metric visual-path-following framework is investigated in different environments. The framework relies on a monocular camera as the only sensing modality. The path is represented as a series of reference images such that each neighboring pair contains a number of common landmarks. Local 3-D geometries are reconstructed between the neighboring reference images to achieve fast feature prediction. This condition allows recovery from tracking failures. During navigation, the robot is controlled using image-based visual servoing. The focus of this paper is on the results from a number of experiments that were conducted in different environments, lighting conditions, and seasons. The experiments with a robot car show that the framework is robust to moving objects and moderate illumination changes. It is also shown that the system is capable of online path learning. Albert Diosi, Sinisa Segvic, Anthony Remazeilles, François Chaumette |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2009 | A mapping and localization framework for scalable appearance-based navigation
Sinisa Segvic, Anthony Remazeilles, Albert Diosi, François Chaumette |
Comput. Vis. Image Underst. | 1 |
| 2008 | Online/Realtime Structure and Motion for General Camera ModelsabstractThis paper presents a novel algorithm for online structure and motion estimation. The algorithm works for general camera models and minimizes object space error, it does not rely on gradient-based optimization, and it is provably globally convergent. In comparison to previous work, which reports cubic complexity in the number of frames, our major contribution is a significant reduction of complexity. The new algorithm requires constant time per frame and can thus be used in online applications. Experimental results show high reconstruction accuracy with respect to simulated ground truth data. We also present two applications in artificial marker reconstruction and handheld augmented reality. Gerald Schweighofer, Sinisa Segvic, Axel Pinz |
WACV | 2 |
| 2007 | A Framework for Scalable Vision-Only Navigation
Sinisa Segvic, Anthony Remazeilles, Albert Diosi, François Chaumette |
ACIVS | 1 |
| 2007 | Large scale vision-based navigation without an accurate global reconstructionabstractAutonomous cars will likely play an important role in the future. A vision system designed to support outdoor navigation for such vehicles has to deal with large dynamic environments, changing imaging conditions, and temporary occlusions by other moving objects. This paper presents a novel appearance-based navigation framework relying on a single perspective vision sensor, which is aimed towards resolving of the above issues. The solution is based on a hierarchical environment representation created during a teaching stage, when the robot is controlled by a human operator. At the top level, the representation contains a graph of key-images with extracted 2D features enabling a robust navigation by visual servoing. The information stored at the bottom level enables to efficiently predict the locations of the features which are currently not visible, and eventually (re-)start their tracking. The outstanding property of the proposed framework is that it enables robust and scalable navigation without requiring a globally consistent map, even in interconnected environments. This result has been confirmed by realistic off-line experiments and successful real-time navigation trials in public urban areas. Sinisa Segvic, Anthony Remazeilles, Albert Diosi, François Chaumette |
CVPR | 1 |
| 2007 | Influence of numerical conditioning on the accuracy of relative orientationabstractWe study the influence of numerical conditioning on the accuracy of two closed-form solutions to the overconstrained relative orientation problem. We consider the well known eight-point algorithm and the recent five-point algorithm, and evaluate changes in their performance due to Hartley's normalization and Muehlich's equilibration. The need for numerical conditioning is introduced by explaining the known occurence of the bias of the eight-point algorithm towards the forward motion. Then it is shown how conditioning can be used to improve the results of the recent five-point algorithm. This is not straightforward since the conditioning disturbs the calibration of the input data. The conditioning therefore needs to be reverted before enforcing the internal cubic constraints of the essential matrix. The obtained improvements are less dramatic than in the case of the eight-point algorithm, for which we offer a plausible explanation. The theoretical claims are backed up with extensive experimentation on noisy artificial datasets, under a variety of geometric and imaging parameters. Sinisa Segvic, Gerald Schweighofer, Axel Pinz |
CVPR | 1 |
| 2007 | Outdoor visual path following experimentsabstractIn this paper the performance of a topological- metric visual path following framework is investigated in different environments. The framework relies on a monocular camera as the only sensing modality. The path is represented as a series of reference images such that each neighboring pair contains a number of common landmarks. Local 3D geometries are reconstructed between the neighboring reference images in order to achieve fast feature prediction which allows the recovery from tracking failures. During navigation the robot is controlled using image-based visual servoing. The experiments show that the framework is robust against moving objects and moderate illumination changes. It is also shown that the system is capable of on-line path learning. Albert Diosi, Anthony Remazeilles, Sinisa Segvic, François Chaumette |
IROS | 3 |
| 2007 | Visual path following using only monocular vision for urban environmentsabstractThis document provides a summary to a short video with the same title. The video shows the French intelligent transportation vehicle CyCab performing visual path following using only monocular vision. All phases of the process are shown with a spoken commentary. In the teaching phase, the user drives the robot manually while images from the camera are stored. Key images with corresponding images features are stored as a map together with 2D and 3D local information. In the navigation phase, CyCab follows the learned path by tracking the images features projected from the map and with a simple visual servoing control law. Albert Diosi, Fabien Spindler, Anthony Remazeilles, Sinisa Segvic, François Chaumette |
IROS | 4 |
| 2006 | Enhancing the Point Feature Tracker by Adaptive Modelling of the Feature Support
Sinisa Segvic, Anthony Remazeilles, François Chaumette |
ECCV (2) | 1 |
| 2003 | A Software Architecture for Distributed Visual Tracking in a Global Vision Localization System
Sinisa Segvic, Slobodan Ribaric |
ICVS | 1 |