Tomás Svoboda

dblp:23/5289 · DBLP profile ↗
← Back
26ranked-venue papers
4as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-authorSystems, architecture and hardware · 5 · 3 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Manual, Semi or Fully Autonomous Flipper Control? A Framework for Fair Comparison
abstract
We investigated the performance of existing semiand fully autonomous methods for controlling flipper-based skid-steer robots. Our study involves the reimplementation of these methods for a fair comparison, and it introduces a novel semi-autonomous control policy that provides a compelling trade-off among current state-of-the-art approaches. We also propose new metrics for assessing cognitive load and traversal quality and offer a benchmarking interface for generating Quality-Load graphs from recorded data. Our results, presented in a 2D Quality-Load space, demonstrate that the new control policy effectively bridges the gap between autonomous and manual control methods. Additionally, we reveal a surprising fact that fully manual, continuous control of all six degrees of freedom remains highly effective when performed by an experienced operator on a well-designed analog controller from a third-person view.
Valentýn Cíhala, Martin Pecka, Tomás Svoboda, Karel Zimmermann
ICRA3
2024 MonoForce: Self-supervised Learning of Physics-informed Model for Predicting Robot-terrain Interaction
abstract
While autonomous navigation of mobile robots on rigid terrain is a well-explored problem, navigating on deformable terrain such as tall grass or bushes remains a challenge. To address it, we introduce an explainable, physics-aware and end-to-end differentiable model which predicts the outcome of robot-terrain interaction from camera images, both on rigid and non-rigid terrain. The proposed MonoForce model consists of a black-box module which predicts robot-terrain interaction forces from onboard cameras, followed by a white-box module, which transforms these forces and a control signals into predicted trajectories, using only the laws of classical mechanics. The differentiable white-box module allows backpropagating the predicted trajectory errors into the black-box module, serving as a self-supervised loss that measures consistency between the predicted forces and ground-truth trajectories of the robot. Experimental evaluation on a public dataset and our data has shown that while the prediction capabilities are comparable to state-of-the-art algorithms on rigid terrain, MonoForce shows superior accuracy on nonrigid terrain such as tall grass or bushes. To facilitate the reproducibility of our results, we release both the code and datasets.
Ruslan Agishev, Karel Zimmermann, Vladimir Kubelka, Martin Pecka, Tomás Svoboda
IROS5
2024 PreCNet: Next-Frame Video Prediction Based on Predictive Coding
abstract
Predictive coding, currently a highly influential theory in neuroscience, has not been widely adopted in machine learning yet. In this work, we transform the seminal model of Rao and Ballard (1999) into a modern deep learning framework while remaining maximally faithful to the original schema. The resulting network we propose (PreCNet) is tested on a widely used next-frame video prediction benchmark, which consists of images from an urban environment recorded from a car-mounted camera, and achieves state-of-the-art performance. Performance on all measures (MSE, PSNR, and SSIM) was further improved when a larger training set (2M images from BDD100k) pointed to the limitations of the KITTI training set. This work demonstrates that an architecture carefully based on a neuroscience model, without being explicitly tailored to the task at hand, can exhibit exceptional performance.
Zdenek Straka, Tomás Svoboda, Matej Hoffmann
IEEE Trans. Neural Networks Learn. Syst.2
2023 T-UDA: Temporal Unsupervised Domain Adaptation in Sequential Point Clouds
abstract
Deep perception models have to reliably cope with an open-world setting of domain shifts induced by different geographic regions, sensor properties, mounting positions, and several other reasons. Since covering all domains with annotated data is technically intractable due to the endless possible variations, researchers focus on unsupervised domain adaptation (UDA) methods that adapt models trained on one (source) domain with annotations available to another (target) domain for which only unannotated data are available. Current predominant methods either leverage semi-supervised approaches, e.g., teacher-student setup, or exploit privileged data, such as other sensor modalities or temporal data consistency. We introduce a novel domain adaptation method that leverages the best of both approaches. Our approach combines input data's temporal and cross-sensor geometric consistency with the mean teacher method. Dubbed T-UDA for “temporal UDA”, such a combination yields massive performance gains for the task of 3D semantic segmentation of driving scenes. Experiments are conducted on Waymo Open Dataset, nuScenes, and SemanticKITTI, for two popular 3D point cloud architectures, Cylinder3D and MinkowskiNet. Our codes are publicly available on https://github.com/ctu-vras/T-UDA.
Awet Haileslassie Gebrehiwot, David Hurych, Karel Zimmermann, Patrick Pérez, Tomás Svoboda
IROS5
2022 Security Consideration of BIA Utilization in Smart Electricity Metering Systems
Vladimir Sobeslav, Josef Horalek, Tomás Svoboda, Hana Dubravova
ICCCI3
2022 Additive Types in Quantitative Type Theory
Vít Sefl, Tomás Svoboda
WoLLIC2
2022 Learning to Predict Lidar Intensities
abstract
We propose a data-driven method for simulating lidar sensors. The method reads computer-generated data, and (i) extracts geometrically simulated lidar point clouds and (ii) predicts the strength of the lidar response –lidar intensities. Qualitative evaluation of the proposed pipeline demonstrates the ability to predict systematic failures such as no/low responses on polished parts of car bodyworks and windows, or strong responses on reflective surfaces such as traffic signs and license/registration plates. We also experimentally show that enhancing the training set by such simulated data improves the segmentation accuracy on the real dataset with limited access to real data. Implementation of the resulting lidar simulator for the GTA V game, as well as the accompanying large dataset, is made publicly available.
Patrik Vacek, Otakar Jasek, Karel Zimmermann, Tomás Svoboda
IEEE Trans. Intell. Transp. Syst.4
2017 Learning for Active 3D Mapping
abstract
We propose an active 3D mapping method for depth sensors, which allow individual control of depth-measuring rays, such as the newly emerging solid-state lidars. The method simultaneously (i) learns to reconstruct a dense 3D occupancy map from sparse depth measurements, and (ii) optimizes the reactive control of depth-measuring rays. To make the first step towards the online control optimization, we propose a fast prioritized greedy algorithm, which needs to update its cost function in only a small fraction of possible rays. The approximation ratio of the greedy algorithm is derived. An experimental evaluation on the subset of the KITTI dataset demonstrates significant improvement in the 3D map accuracy when learning-to-reconstruct from sparse measurements is coupled with the optimization of depth measuring rays.
Karel Zimmermann, Tomás Petrícek 0002, Vojtech Salanský, Tomás Svoboda
ICCV4
2017 Fast simulation of vehicles with non-deformable tracks
abstract
This paper presents a novel technique that allows for both computationally fast and sufficiently plausible simulation of vehicles with non-deformable tracks. The method is based on an effect we have called Contact Surface Motion. A comparison with several other methods for simulation of tracked vehicle dynamics is presented with the aim to evaluate methods that are available off-the-shelf or with minimum effort in general-purpose robotics simulators. The proposed method is implemented as a plugin for the open-source physics-based simulator Gazebo using the Open Dynamics Engine.
Martin Pecka, Karel Zimmermann, Tomás Svoboda
IROS3
2016 Design Solution of Centralized Monitoring System of Airport Facilities
Josef Horalek, Tomás Svoboda
ICCCI (2)2
2016 Autonomous flipper control with safety constraints
abstract
Policy Gradient methods require many real-world trials. Some of the trials may endanger the robot system and cause its rapid wear. Therefore, a safe or at least gentle-to-wear exploration is a desired property. We incorporate bounds on the probability of unwanted trials into the recent Contextual Relative Entropy Policy Search method. The proposed algorithm is evaluated on the task of autonomous flipper control for a real Search and Rescue rover platform.
Martin Pecka, Vojtech Salanský, Karel Zimmermann, Tomás Svoboda
IROS4
2015 Analysis of the Use of Cloud Services and Their Effects on the Efficient Functioning of a Company
Josef Horalek, Simeon Karamazov, Filip Holík, Tomás Svoboda
ICCCI (2)4
2014 Non-Rigid Object Detection with LocalInterleaved Sequential Alignment (LISA)
abstract
This paper shows that the successively evaluated features used in a sliding window detection process to decide about object presence/absence also contain knowledge about object deformation. We exploit these detection features to estimate the object deformation. Estimated deformation is then immediately applied to not yet evaluated features to align them with the observed image data. In our approach, the alignment estimators are jointly learned with the detector. The joint process allows for the learning of each detection stage from less deformed training samples than in the previous stage. For the alignment estimation we propose regressors that approximate non-linear regression functions and compute the alignment parameters extremely fast.
Karel Zimmermann, David Hurych, Tomás Svoboda
IEEE Trans. Pattern Anal. Mach. Intell.3
2012 Exploiting Features - Locally Interleaved Sequential Alignment for Object Detection
Karel Zimmermann, David Hurych, Tomás Svoboda
ACCV (1)3
2012 Area-weighted surface normals for 3D object recognition
Tomás Petrícek 0002, Tomás Svoboda
ICPR2
2009 Anytime learning for the NoSLLiP tracker
Karel Zimmermann, Tomás Svoboda, Jiri Matas
Image Vis. Comput.2
2009 Tracking by an Optimal Sequence of Linear Predictors
abstract
We propose a learning approach to tracking explicitly minimizing the computational complexity of the tracking process subject to user-defined probability of failure (loss-of-lock) and precision. The tracker is formed by a Number of Sequences of Learned Linear Predictors (NoSLLiP). Robustness of NoSLLiP is achieved by modeling the object as a collection of local motion predictors--object motion is estimated by the outlier-tolerant RANSAC algorithm from local predictions. Efficiency of the NoSLLiP tracker stems from (i) the simplicity of the local predictors and (ii) from the fact that all design decisions--the number of local predictors used by the tracker, their computational complexity (i.e. the number of observations the prediction is based on), locations as well as the number of RANSAC iterations are all subject to the optimization (learning) process. All time-consuming operations are performed during the learning stage--tracking is reduced to only a few hundreds integer multiplications in each step. On PC with 1xK8 3200+, a predictor evaluation requires about 30 microseconds. The proposed approach is verified on publicly-available sequences with approximately 12000 frames with ground-truth. Experiments demonstrates, superiority in frame rates and robustness with respect to the SIFT detector, Lucas-Kanade tracker and other trackers.
Karel Zimmermann, Jiri Matas, Tomás Svoboda
IEEE Trans. Pattern Anal. Mach. Intell.3
2007 Adaptive Parameter Optimization for Real-time Tracking
abstract
Adaptation of a tracking procedure combined in a common way with a Kalman filter is formulated as an constrained optimization problem, where a trade-off between precision and loss-of-lock probability is explicitly taken into account. While the tracker is learned in order to minimize computational complexity during a learning stage, in a tracking stage the precision is maximized online under a constraint imposed by the loss-of-lock probability resulting in an optimal setting of the tracking procedure. We experimentally show that the proposed method converges to a steady solution in all variables. In contrast to a common Kalman filter based tracking, we achieve a significantly lower state covariance matrix. We also show, that if the covariance matrix is continuously updated, the method is able to adapt to a different situations. If a dynamic model is precise enough the tracker is allowed to spend a longer time with a fine motion estimation, however, if the motion gets saccadic, i.e. unpredictable by the dynamic model, the method automatically gives up the precision in order to avoid loss-of-lock.
Karel Zimmermann, Tomás Svoboda, Jiri Matas
ICCV2
2006 Special issue: Omnidirectional vision and camera networks
Peter F. Sturm, Tomás Svoboda, Seth J. Teller
Comput. Vis. Image Underst.2
2003 Fast indexing for image retrieval based on local appearance with re-ranking
abstract
This paper describes an approach to retrieve images containing specific objects, scenes or buildings. The image content is captured by a set of local features. More precisely, we use so-called invariant regions. These are features with shapes that self-adapt to the viewpoint. The physical parts on the object surface that they carve out are the same in all views, even though the extraction proceeds from a single view only. The surface patterns within the regions are then characterized by a feature vector of moment invariants. Invariance is under affine geometric deformations and scaled color bands with an offset added. This allows regions from different views to be matched efficiently. An indexing technique based on vantage point tree organizes the feature vectors in such a way that a naive sequential search can be avoided. This results in sublinear computation times to retrieve images from a database. In order to get sufficient certainty about the correctness of the retrieved images, a method to increase the number of matched regions is introduced. This way, the system is both efficient and discriminant. It is demonstrated how scenes or buildings are recognized, even in case of partial visibility and under a large variety of viewing condition changes.
Hao Shao, Tomás Svoboda, Vittorio Ferrari, Tinne Tuytelaars, Luc Van Gool
ICIP (3)2
2003 Monkeys -- A Software Architecture for ViRoom -- Low-Cost Multicamera System
Petr Doubek, Tomás Svoboda, Luc Van Gool
ICVS2
2003 blue-c: a spatially immersive display and 3D video portal for telepresence
abstract
We present blue-c , a new immersive projection and 3D video acquisition environment for virtual design and collaboration. It combines simultaneous acquisition of multiple live video streams with advanced 3D projection technology in a CAVE™-like environment, creating the impression of total immersion. The blue-c portal currently consists of three rectangular projection screens that are built from glass panels containing liquid crystal layers. These screens can be switched from a whitish opaque state (for projection) to a transparent state (for acquisition), which allows the video cameras to "look through" the walls. Our projection technology is based on active stereo using two LCD projectors per screen. The projectors are synchronously shuttered along with the screens, the stereo glasses, active illumination devices, and the acquisition hardware. From multiple video streams, we compute a 3D video representation of the user in real time. The resulting video inlays are integrated into a networked virtual environment. Our design is highly scalable, enabling blue-c to connect to portals with less sophisticated hardware.
Markus Gross 0001, Stephan Würmlin, Martin Näf, Edouard Lamboray, Christian P. Spagno, Andreas M. Kunz, Esther Koller-Meier, Tomás Svoboda, Luc Van Gool, Silke Lang, Kai Strehlke, Andrew Vande Moere, Oliver G. Staadt
ACM Trans. Graph.8
2002 Epipolar Geometry for Central Catadioptric Cameras
Tomás Svoboda, Tomás Pajdla
Int. J. Comput. Vis.1
2001 Matching in Catadioptric Images with Appropriate Windows, and Outliers Removal
Tomás Svoboda, Tomás Pajdla
CAIP1
1998 Epipolar Geometry of Panoramic Cameras
Tomás Svoboda, Tomás Pajdla, Václav Hlavác
ECCV (1)1
1997 A Badly Calibrated Camera in Ego-Motion Estimation Propagation of Uncertainty
Tomás Svoboda, Peter F. Sturm
CAIP1