EDBT 2026 Demo / reviewers in the wild / expert
Vladimír Petrík
dblp:153/7849
· DBLP profile ↗
12ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0001-9450-4987ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 5 since 2021Systems, architecture and hardware · 8 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | 6D Object Pose Tracking in Internet Videos for Robotic ManipulationabstractWe seek to extract a temporally consistent 6D pose trajectory of a manipulated object from an Internet instructional video. This is a challenging set-up for current 6D pose estimation methods due to uncontrolled capturing conditions, subtle but dynamic object motions, and the fact that the exact mesh of the manipulated object is not known. To address these challenges, we present the following contributions. First, we develop a new method that estimates the 6D pose of any object in the input image without prior knowledge of the object itself. The method proceeds by (i) retrieving a CAD model similar to the depicted object from a large-scale model database, (ii) 6D aligning the retrieved CAD model with the input image, and (iii) grounding the absolute scale of the object with respect to the scene. Second, we extract smooth 6D object trajectories from Internet videos by carefully tracking the detected objects across video frames. The extracted object trajectories are then retargeted via trajectory optimization into the configuration space of a robotic manipulator. Third, we thoroughly evaluate and ablate our 6D pose estimation method on YCB-V and HOPE-Video datasets as well as a new dataset of instructional videos manually annotated with approximate 6D object trajectories. We demonstrate significant improvements over existing state-of-the-art RGB 6D pose estimation methods. Finally, we show that the 6D object motion estimated from Internet videos can be transferred to a 7-axis robotic manipulator both in a virtual simulator as well as in a real world set-up. We also successfully apply our method to egocentric videos taken from the EPIC-KITCHENS dataset, demonstrating potential for Embodied AI applications. Georgy Ponimatkin, Martin Cífka, Tomás Soucek, Médéric Fourmy, Yann Labbé, Vladimír Petrík, Josef Sivic |
ICLR | 6 |
| 2025 | FocalPose++: Focal Length and Object Pose Estimation via Render and CompareabstractWe introduce FocalPose++, a neural render-and-compare method for jointly estimating the camera-object 6D pose and camera focal length given a single RGB input image depicting a known object. The contributions of this work are threefold. First, we derive a focal length update rule that extends an existing state-of-the-art render-and-compare 6D pose estimator to address the joint estimation task. Second, we investigate several different loss functions for jointly estimating the object pose and focal length. We find that a combination of direct focal length regression with a reprojection loss disentangling the contribution of translation, rotation, and focal length leads to improved results. Third, we explore the effect of different synthetic training data on the performance of our method. Specifically, we investigate different distributions used for sampling object's 6D pose and camera's focal length when rendering the synthetic images, and show that parametric distribution fitted on real training data works the best. We show results on three challenging benchmark datasets that depict known 3D models in uncontrolled settings. We demonstrate that our focal length and 6D pose estimates have lower error than the existing state-of-the-art methods. Martin Cífka, Georgy Ponimatkin, Yann Labbé, Bryan C. Russell, Mathieu Aubry, Vladimír Petrík, Josef Sivic |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | GJK++: Leveraging Acceleration Methods for Faster Collision DetectionabstractCollision detection is a fundamental problem in various domains, such as robotics, computational physics, and computer graphics. In general, collision detection is tackled as a computational geometry problem, with the so-called Gilbert, Johnson, and Keerthi (GJK) algorithm being the most adopted solution nowadays. While introduced in 1988, GJK remains the most effective solution to compute the distance or the collision between two 3D convex geometries. Over the years, it was shown to be efficient, scalable, and generic, operating on a broad class of convex shapes, ranging from simple primitives (sphere, ellipsoid, box, cone, capsule, etc.) to complex meshes involving thousands of vertices. In this article, we introduce several contributions to accelerate collision detection and distance computation between convex geometries by leveraging the fact that these two problems are fundamentally optimization problems. Notably, we establish that the GJK algorithm is a specific sub-case of the well-established Frank-Wolfe (FW) algorithm in convex optimization. By adapting recent works linking Polyak and Nesterov accelerations to Frank-Wolfe methods, we also propose two accelerated extensions of the classic GJK algorithm. Through an extensive benchmark over millions of collision pairs involving objects of daily life, we show that these two accelerated GJK extensions significantly reduce the overall computational burden of collision detection, leading to computation times that are up to two times faster. Finally, we hope this work will significantly reduce the computational cost of modern robotic simulators, allowing the speed-up of modern robotic applications that heavily rely on simulation, such as reinforcement learning or trajectory optimization. Louis Montaut, Quentin Le Lidec, Vladimír Petrík, Josef Sivic, Justin Carpentier |
IEEE Trans. Robotics | 3 |
| 2023 | Differentiable Collision Detection: a Randomized Smoothing ApproachabstractCollision detection is an important component of many robotics applications, from robot control to simulation, including motion planning and estimation. While the seminal works on the topic date back to the 80s, it is only recently that the question of properly differentiating collision detection has emerged as a central issue, thanks notably to the ongoing and various efforts made by the scientific community around the topic of differentiable physics. Yet, very few solutions have been suggested so far, and only with a strong assumption on the nature of the shapes involved. In this work, we introduce a generic and efficient approach to compute the derivatives of collision detection for any pair of convex shapes, by notably leveraging randomized smoothing techniques which have shown to be particularly adapted to capture the derivatives of non-smooth problems. This approach is implemented in the HPP-FCL and Pinocchio ecosystems, and evaluated on classic datasets and problems of the robotics literature, demonstrating few micro-second timings to compute informative derivatives directly exploitable by many real robotic applications, including differentiable simulation. Louis Montaut, Quentin Le Lidec, Antoine Bambade, Vladimír Petrík, Josef Sivic, Justin Carpentier |
ICRA | 4 |
| 2023 | Multi-Contact Task and Motion Planning Guided by Video DemonstrationabstractThis work aims at leveraging instructional video to guide the solving of complex multi-contact task-and-motion planning tasks in robotics. Towards this goal, we propose an extension of the well-established Rapidly-Exploring Random Tree (RRT) planner, which simultaneously grows multiple trees around grasp and release states extracted from the guiding video. Our key novelty lies in combining contact states, and 3D object poses extracted from the guiding video with a traditional planning algorithm that allows us to solve tasks with sequential dependencies, for example, if an object needs to be placed at a specific location to be grasped later. To demonstrate the benefits of the proposed video-guided planning approach, we design a new benchmark with three challenging tasks: (i) 3D re-arrangement of multiple objects between a table and a shelf, (ii) multi-contact transfer of an object through a tunnel, and (iii) transferring objects using a tray in a similar way a waiter transfers dishes. We demonstrate the effectiveness of our planning algorithm on several robots, including the Franka Emika Panda and the KUKA KMR iiwa. Kateryna Zorina, David Kovár, Florent Lamiraux, Nicolas Mansard, Justin Carpentier, Josef Sivic, Vladimír Petrík |
ICRA | 7 |
| 2022 | Learning Object Manipulation Skills from Video via Approximate Differentiable PhysicsabstractWe aim to teach robots to perform simple object manipulation tasks by watching a single video demonstration. Towards this goal, we propose an optimization approach that outputs a coarse and temporally evolving 3D scene to mimic the action demonstrated in the input video. Similar to previous work, a differentiable renderer ensures perceptual fidelity between the 3D scene and the 2D video. Our key novelty lies in the inclusion of a differentiable approach to solve a set of Ordinary Differential Equations (ODEs) that allows us to approximately model laws of physics such as gravity, friction, and hand-object or object-object interactions. This not only enables us to dramatically improve the quality of estimated hand and object states, but also produces physically admissible trajectories that can be directly translated to a robot without the need for costly reinforcement learning. We evaluate our approach on a 3D reconstruction task that consists of 54 video demonstrations sourced from 9 actions such as pull something from right to left or put something in front of something. Our approach improves over previous state-of-the-art by almost 30%, demonstrating superior quality on especially challenging actions involving physical interactions of two objects such as put something onto something. Finally, we showcase the learned skills on a Franka Emika Panda robot. Vladimír Petrík, Mohammad Nomaan Qureshi, Josef Sivic, Makarand Tapaswi |
IROS | 1 |
| 2019 | Feedback-based Fabric Strip FoldingabstractAccurate manipulation of a deformable body such as a piece of fabric is difficult because of its many degrees of freedom and unobservable properties affecting its dynamics. To alleviate these challenges, we propose the application of feedback-based control to robotic fabric strip folding. The feedback is computed from the low dimensional state extracted from a camera image. We trained the controller using reinforcement learning in simulation which was calibrated to cover the real fabric strip behaviors. The proposed feedback-based folding was experimentally compared to two state-of-the-art folding methods and our method outperformed both of them in terms of accuracy. Vladimír Petrík, Ville Kyrki |
IROS | 1 |
| 2018 | Automatic Material Properties Estimation for the Physics-Based Robotic Garment FoldingabstractThe estimation of the fabric material property during the folding is presented. The available techniques for the accurate garment folding rely on known material properties. Currently, the properties are estimated by an operator in advance of folding. We propose an iterative strategy, which updates the property while the garment is folded. The estimation is formulated as an optimisation task. It is based on measurements from a laser range finder. The proposed algorithm improves the estimation iteratively and prevents the garment from slipping at the same time. We demonstrate the estimation procedure for 10 fabric strips of different materials. Vladimír Petrík, Jakub Cmiral, Vladimír Smutný, Pavel Krsek, Václav Hlavác |
ICRA | 1 |
| 2017 | Model-free approach to garments unfolding based on detection of folded layersabstractThe proposed work deals with robotic unfolding of a garment that has been placed flat on a table and folded over a certain axis. The algorithm combines image and depth data to detect the bottom and top (folded) layer of the garment. The detection is formulated as a labeling of the garment surface and solved in an energy minimization framework. Once the garment pose is known, several candidate folding axes are generated and used to unfold the garment virtually. The correct folding axis is selected from these candidate axes. The method does not set any constraints on the garment shape; thus it can deal with various types of garments including jackets, pants, shorts, skirts or T-shirts of any sleeve lengths. The garment is unfolded by the dual-arm robot. One arm grasps boundary of the top layer and brings it over the estimated folding axis, while the second arm is holding the bottom layer to prevent the garment from slipping. The perception procedure was tested on the annotated dataset that we are making publicly available. The experimental evaluation of the robotic manipulation is also provided. Jan Stria, Vladimír Petrík, Václav Hlavác |
IROS | 2 |
| 2016 | Physics-based model of a rectangular garment for robotic foldingabstractThe ability to perform an accurate robotic fold is essential to obtain the properly folded garment. Available solutions rely on a rough folding surface or on a comprehensive simulation, both preventing the garment from slipping on the table during folding. This paper proposes a new algorithm for a folding path design respecting the garment material properties and preventing the garment slipping. The folding path is derived based on the equilibrium of forces under the simplifying assumptions of a rectangular and homogeneous garment. This approach allows folding the rectangular garment on a low friction table surface as we demonstrated in the experiments performed by a dual-arm robotic testbed. Vladimír Petrík, Vladimír Smutný, Pavel Krsek, Václav Hlavác |
IROS | 1 |
| 2016 | Folding Clothes Autonomously: A Complete PipelineabstractThis work presents a complete pipeline for folding a pile of clothes using a dual-armed robot. This is a challenging task both from the viewpoint of machine vision and robotic manipulation. The presented pipeline is comprised of the following parts: isolating and picking up a single garment from a pile of crumpled garments, recognizing its category, unfolding the garment using a series of manipulations performed in the air, placing the garment roughly flat on a work table, spreading it, and, finally, folding it in several steps. The pile is segmented into separate garments using color and texture information, and the ideal grasping point is selected based on the features computed from a depth map. The recognition and unfolding of the hanging garment are performed in an active manner, utilizing the framework of active random forests to detect grasp points, while optimizing the robot actions. The spreading procedure is based on the detection of deformations of the garment's contour. The perception for folding employs fitting of polygonal models to the contour of the observed garment, both spread and already partially folded. We have conducted several experiments on the full pipeline producing very promising results. To our knowledge, this is the first work addressing the complete unfolding and folding pipeline on a variety of garments, including T-shirts, towels, and shorts. Andreas Doumanoglou, Jan Stria, Georgia Peleka, Ioannis Mariolis, Vladimír Petrík, Andreas Kargakos, Libor Wagner, Václav Hlavác, Tae-Kyun Kim 0001, Sotiris Malassiotis |
IEEE Trans. Robotics | 5 |
| 2014 | Garment perception and its folding using a dual-arm robotabstractThe work addresses the problem of clothing perception and manipulation by a two armed industrial robot aiming at a real-time automated folding of a piece of garment spread out on a flat surface. A complete solution combining vision sensing, garment segmentation and understanding, planning of the manipulation and its real execution on a robot is proposed. A new polygonal model of a garment is introduced. Fitting the model into a segmented garment contour is used to detect garment landmark points. It is shown how folded variants of the unfolded model can be derived automatically. Universality and usefulness of the model is demonstrated by its favorable performance within the whole folding procedure which is applicable to a variety of garments categories (towel, pants, shirt, etc.) and evaluated experimentally using the two armed robot. The principal novelty with respect to the state of the art is in the new garment polygonal model and its manipulation planning algorithm which leads to the speed up by two orders of magnitude. Jan Stria, Daniel Prusa, Václav Hlavác, Libor Wagner, Vladimír Petrík, Pavel Krsek, Vladimír Smutný |
IROS | 5 |