Tomás Soucek

dblp:213/5738 · DBLP profile ↗
← Back
21ranked-venue papers
6as first author
13since 2021 · last 2025
0000-0001-6911-5517ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 5 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions
abstract
The goal of this work is to generate step-by-step visual instructions in the form of a sequence of images, given an input image that provides the scene context and the sequence of textual instructions. This is a challenging problem as it requires generating multi-step image sequences to achieve a complex goal while being grounded in a specific environment. Part of the challenge stems from the lack of large-scale training data for this problem. The contribution of this work is thus three-fold. First, we introduce an automatic approach for collecting large step-by-step visual instruction training data from instructional videos. We apply this approach to one million videos and create a large-scale, high-quality dataset of 0.6M sequences of image-text pairs. Second, we develop and train ShowHowTo, a video diffusion model capable of generating step-by-step visual instructions consistent with the provided input image. Third, we evaluate the generated image sequences across three dimensions of accuracy (step, scene, and task) and show our model achieves state-of-the-art results on all of them. Our code, dataset, and trained models are publicly available.
Tomás Soucek, Prajwal Gatti, Michael Wray, Ivan Laptev, Dima Damen, Josef Sivic
CVPR1
2025 6D Object Pose Tracking in Internet Videos for Robotic Manipulation
abstract
We seek to extract a temporally consistent 6D pose trajectory of a manipulated object from an Internet instructional video. This is a challenging set-up for current 6D pose estimation methods due to uncontrolled capturing conditions, subtle but dynamic object motions, and the fact that the exact mesh of the manipulated object is not known. To address these challenges, we present the following contributions. First, we develop a new method that estimates the 6D pose of any object in the input image without prior knowledge of the object itself. The method proceeds by (i) retrieving a CAD model similar to the depicted object from a large-scale model database, (ii) 6D aligning the retrieved CAD model with the input image, and (iii) grounding the absolute scale of the object with respect to the scene. Second, we extract smooth 6D object trajectories from Internet videos by carefully tracking the detected objects across video frames. The extracted object trajectories are then retargeted via trajectory optimization into the configuration space of a robotic manipulator. Third, we thoroughly evaluate and ablate our 6D pose estimation method on YCB-V and HOPE-Video datasets as well as a new dataset of instructional videos manually annotated with approximate 6D object trajectories. We demonstrate significant improvements over existing state-of-the-art RGB 6D pose estimation methods. Finally, we show that the 6D object motion estimated from Internet videos can be transferred to a 7-axis robotic manipulator both in a virtual simulator as well as in a real world set-up. We also successfully apply our method to egocentric videos taken from the EPIC-KITCHENS dataset, demonstrating potential for Embodied AI applications.
Georgy Ponimatkin, Martin Cífka, Tomás Soucek, Médéric Fourmy, Yann Labbé, Vladimír Petrík, Josef Sivic
ICLR3
2025 Watermarking Autoregressive Image Generation
abstract
Watermarking the outputs of generative models has emerged as a promising approach for tracking their provenance. Despite significant interest in autoregressive image generation models and their potential for misuse, no prior work has attempted to watermark their outputs at the token level. In this work, we present the first such approach by adapting language model watermarking techniques to this setting. We identify a key challenge: the lack of reverse cycle-consistency (RCC), wherein re-tokenizing generated image tokens significantly alters the token sequence, effectively erasing the watermark. To address this and to make our method robust to common image transformations, neural compression, and removal attacks, we introduce (i) a custom tokenizer-detokenizer finetuning procedure that improves RCC, and (ii) a complementary watermark synchronization layer. As our experiments demonstrate, our approach enables reliable and robust watermark detection with theoretically grounded p-values. Code and models are available at https://github.com/facebookresearch/wmar.
Nikola Jovanovic 0001, Ismail Labiad, Tomás Soucek, Martin T. Vechev, Pierre Fernandez
NeurIPS3
2025 Transferable Black-Box One-Shot Forging of Watermarks via Image Preference Models
abstract
Recent years have seen a surge in interest in digital content watermarking techniques, driven by the proliferation of generative models and increased legal pressure. With an ever-growing percentage of AI-generated content available online, watermarking plays an increasingly important role in ensuring content authenticity and attribution at scale. There have been many works assessing the robustness of watermarking to removal attacks, yet, watermark forging, the scenario when a watermark is stolen from genuine content and applied to malicious content, remains underexplored. In this work, we investigate watermark forging in the context of widely used post-hoc image watermarking. Our contributions are as follows. First, we introduce a preference model to assess whether an image is watermarked. The model is trained using a ranking loss on purely procedurally generated images without any need for real watermarks. Second, we demonstrate the model's capability to remove and forge watermarks by optimizing the input image through backpropagation. This technique requires only a single watermarked image and works without knowledge of the watermarking model, making our attack much simpler and more practical than attacks introduced in related work. Third, we evaluate our proposed method on a variety of post-hoc image watermarking models, demonstrating that our approach can effectively forge watermarks, questioning the security of current watermarking approaches. Our code and further resources are publicly available.
Tomás Soucek, Sylvestre-Alvise Rebuffi, Pierre Fernandez, Nikola Jovanovic 0001, Hady ElSahar, Valeriu Lacatusu, Alexandre Mourachko
NeurIPS1
2024 GenHowTo: Learning to Generate Actions and State Transformations from Instructional Videos
abstract
We address the task of generating temporally consistent and physically plausible images of actions and object state transformations. Given an input image and a text prompt describing the targeted transformation, our generated images preserve the environment and transform objects in the initial image. Our contributions are threefold. First, we leverage a large body of instructional videos and automatically mine a dataset of triplets of consecutive frames corresponding to initial object states, actions, and resulting object transformations. Second, equipped with this data, we develop and train a conditioned diffusion model dubbed GenHowTo. Third, we evaluate GenHowTo on a variety of objects and actions and show superior performance compared to existing methods. In particular, we introduce a quantitative evaluation where GenHowTo achieves 88% and 74% on seen and unseen interaction categories, respectively, outperforming prior work by a large margin.
Tomás Soucek, Dima Damen, Michael Wray, Ivan Laptev, Josef Sivic
CVPR1
2024 TransNet V2: An Effective Deep Network Architecture for Fast Shot Transition Detection
abstract
Although automatic shot transition detection approaches are already investigated for more than two decades, an effective universal human-level model was not proposed yet. Even for common shot transitions like hard cuts or simple gradual changes, the potential diversity of analyzed video contents may still lead to both false hits and false dismissals. Recently, deep learning-based approaches significantly improved the accuracy of shot transition detection using 3D convolutional architectures and artificially created training data. Nevertheless, one hundred percent accuracy is still an unreachable ideal. In this paper, we share the current version of our deep network TransNet V2 that reaches state-of-the-art performance on respected benchmarks. A trained instance of the model is provided so it can be instantly utilized by the community for a highly efficient analysis of large video archives. Furthermore, the network architecture, as well as our experience with the training process, are detailed, including simple code snippets for convenient usage of the proposed model and visualization of results.
Tomás Soucek, Jakub Lokoc
ACM Multimedia1
2024 Multi-Task Learning of Object States and State-Modifying Actions From Web Videos
abstract
We aim to learn to temporally localize object state changes and the corresponding state-modifying actions by observing people interacting with objects in long uncurated web videos. We introduce three principal contributions. First, we develop a self-supervised model for jointly learning state-modifying actions together with the corresponding object states from an uncurated set of videos from the Internet. The model is self-supervised by the causal ordering signal, i.e., initial object state → manipulating action → end state. Second, we explore alternative multi-task network architectures and identify a model that enables efficient joint learning of multiple object states and actions, such as pouring water and pouring coffee, together. Third, we collect a new dataset, named ChangeIt, with more than 2600 hours of video and 34 thousand changes of object states. We report results on an existing instructional video dataset COIN as well as our new large-scale ChangeIt dataset containing tens of thousands of long uncurated web videos depicting various interactions such as hole drilling, cream whisking, or paper plane folding. We show that our multi-task model achieves a relative improvement of 40% over the prior methods and significantly outperforms both image-based and video-based zero-shot models for this problem.
Tomás Soucek, Jean-Baptiste Alayrac, Antoine Miech, Ivan Laptev, Josef Sivic
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Look for the Change: Learning Object States and State-Modifying Actions from Untrimmed Web Videos
abstract
Human actions often induce changes of object states such as “cutting an apple”, “cleaning shoes” or “pouring coffee”. In this paper, we seek to temporally localize object states (e.g. “empty” and “full” cup) together with the corresponding state-modifying actions (“pouring coffee”) in long uncurated videos with minimal supervision. The contributions of this work are threefold. First, we develop a self-supervised model for jointly learning state-modifying actions together with the corresponding object states from an uncurated set of videos from the Internet. The model is self-supervised by the causal ordering signal, i.e. initialobject state → manipulating action → end state. Second, to cope with noisy uncurated training data, our model incorporates a noise adaptive weighting module supervised by a small number of annotated still images, that allows to efficiently filter out irrelevant videos during training. Third, we collect a new dataset with more than 2600 hours of video and 34 thousand changes of object states, and manually annotate a part of this data to validate our approach. Our results demonstrate substantial improvements over prior work in both action and object state-recognition in video.
Tomás Soucek, Jean-Baptiste Alayrac, Antoine Miech, Ivan Laptev, Josef Sivic
CVPR1
2022 Video Search with Context-Aware Ranker and Relevance Feedback
Jakub Lokoc, Frantisek Mejzlík, Tomás Soucek, Patrik Dokoupil, Ladislav Peska
MMM (2)3
2021 W2VV++ BERT Model at VBS 2021
Ladislav Peska, Gregor Kovalcík, Tomás Soucek, Vít Skrhák, Jakub Lokoc
MMM (2)3
2021 How Many Neighbours for Known-Item Search?
Jakub Lokoc, Tomás Soucek
SISAP2
2021 Interactive Video Retrieval in the Age of Deep Learning - Detailed Evaluation of VBS 2019
abstract
Despite the fact that automatic content analysis has made remarkable progress over the last decade - mainly due to significant advances in machine learning - interactive video retrieval is still a very challenging problem, with an increasing relevance in practical applications. The Video Browser Showdown (VBS) is an annual evaluation competition that pushes the limits of interactive video retrieval with state-of-the-art tools, tasks, data, and evaluation metrics. In this paper, we analyse the results and outcome of the 8th iteration of the VBS in detail. We first give an overview of the novel and considerably larger V3C1 dataset and the tasks that were performed during VBS 2019. We then go on to describe the search systems of the six international teams in terms of features and performance. And finally, we perform an in-depth analysis of the per-team success ratio and relate this to the search strategies that were applied, the most popular features, and problems that were experienced. A large part of this analysis was conducted based on logs that were collected during the competition itself. This analysis gives further insights into the typical search behavior and differences between expert and novice users. Our evaluation shows that textual search and content browsing are the most important aspects in terms of logged user interactions. Furthermore, we observe a trend towards deep learning based features, especially in the form of labels generated by artificial neural networks. But nevertheless, for some tasks, very specific content-based search features are still being used. We expect these findings to contribute to future improvements of interactive video search systems.
Luca Rossetto, Ralph Gasser, Jakub Lokoc, Werner Bailer, Klaus Schöffmann, Bernd Münzer, Tomás Soucek, Phuong Anh Nguyen 0002, Paolo Bolettieri, Andreas Leibetseder, Stefanos Vrochidis
IEEE Trans. Multim.7
2021 Is the Reign of Interactive Search Eternal? Findings from the Video Browser Showdown 2020
abstract
Comprehensive and fair performance evaluation of information retrieval systems represents an essential task for the current information age. Whereas Cranfield-based evaluations with benchmark datasets support development of retrieval models, significant evaluation efforts are required also for user-oriented systems that try to boost performance with an interactive search approach. This article presents findings from the 9th Video Browser Showdown, a competition that focuses on a legitimate comparison of interactive search systems designed for challenging known-item search tasks over a large video collection. During previous installments of the competition, the interactive nature of participating systems was a key feature to satisfy known-item search needs, and this article continues to support this hypothesis. Despite the fact that top-performing systems integrate the most recent deep learning models into their retrieval process, interactive searching remains a necessary component of successful strategies for known-item search tasks. Alongside the description of competition settings, evaluated tasks, participating teams, and overall results, this article presents a detailed analysis of query logs collected by the top three performing systems, SOMHunter, VIRET, and vitrivr. The analysis provides a quantitative insight to the observed performance of the systems and constitutes a new baseline methodology for future events. The results reveal that the top two systems mostly relied on temporal queries before a correct frame was identified. An interaction log analysis complements the result log findings and points to the importance of result set and video browsing approaches. Finally, various outlooks are discussed in order to improve the Video Browser Showdown challenge in the future.
Jakub Lokoc, Patrik Veselý, Frantisek Mejzlík, Gregor Kovalcík, Tomás Soucek, Luca Rossetto, Klaus Schöffmann, Werner Bailer, Cathal Gurrin, Loris Sauter, Jaeyub Song, Stefanos Vrochidis, Jiaxin Wu 0001, Björn Þór Jónsson 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2020 Towards Evaluating and Simulating Keyword Queries for Development of Interactive Known-item Search Systems
abstract
Searching for memorized images in large datasets (known-item search) is a challenging task due to a limited effectiveness of retrieval models as well as limited ability of users to formulate suitable queries and choose an appropriate search strategy. A popular option to approach the task is to automatically detect semantic concepts and rely on interactive specification of keywords during the search session. Nonetheless, employed instances of such search models are often set arbitrarily in existing KIS systems as comprehensive evaluations with reals users are time demanding. This paper envisions and investigates an option to simulate keyword queries in a selected "toy'' (yet competitive) keyword search model relying on a deep image classification network. Specifically, two properties of such keyword-based model are experimentally investigated with our known-item search benchmark dataset: which output transformation and ranking models are effective for the utilized classification model and whether there are some options for simulations of keyword queries. In addition to the main objective, the paper inspects also the effect of interactive query reformulations for the considered keyword search model.
Ladislav Peska, Frantisek Mejzlík, Tomás Soucek, Jakub Lokoc
ICMR3
2020 SOMHunter: Lightweight Video Search System with SOM-Guided Relevance Feedback
abstract
In the last decade, the Video Browser Showdown (VBS) became a comparative platform for various interactive video search tools competing in selected video retrieval tasks. However, the participation of new teams with an own, novel tool is prohibitively time-demanding because of the large number and complexity of components required for constructing a video search system from scratch. To partially alleviate this difficulty, we provide an open-source version of the lightweight known-item search system SOMHunter that competed successfully at VBS 2020. The system combines several features for text-based search initialization and browsing of large result sets; in particular a variant of W2VV++ model for text search, temporal queries for targeting sequences of frames, several types of displays including the eponymous self-organizing map view, and a feedback-based approach for maintaining the relevance scores inspired by PICHunter. The minimalistic, easily extensible implementation of SOMHunter should serve as a solid basis for constructing new search systems, thus facilitating easier exploration of new video retrieval ideas.
Miroslav Kratochvíl, Frantisek Mejzlík, Patrik Veselý, Tomás Soucek, Jakub Lokoc
ACM Multimedia4
2020 A W2VV++ Case Study with Automated and Interactive Text-to-Video Retrieval
abstract
As reported by respected evaluation campaigns focusing both on automated and interactive video search approaches, deep learning started to dominate the video retrieval area. However, the results are still not satisfactory for many types of search tasks focusing on high recall. To report on this challenging problem, we present two orthogonal task-based performance studies centered around the state-of-the-art W2VV++ query representation learning model for video retrieval. First, an ablation study is presented to investigate which components of the model are effective in two types of benchmark tasks focusing on high recall. Second, interactive search scenarios from the Video Browser Showdown are analyzed for two winning prototype systems implementing a selected variant of the model and providing additional querying and visualization components. The analysis of collected logs demonstrates that even with the state-of-the-art text search video retrieval model, it is still auspicious to integrate users into the search process for task types, where high recall is essential.
Jakub Lokoc, Tomás Soucek, Patrik Veselý, Frantisek Mejzlík, Jiaqi Ji, Chaoxi Xu, Xirong Li 0001
ACM Multimedia2
2020 VIRET at Video Browser Showdown 2020
Jakub Lokoc, Gregor Kovalcík, Tomás Soucek
MMM (2)3
2019 VIRET: A Video Retrieval Tool for Interactive Known-item Search
abstract
Known-item search in large video collections still represents a challenging task for current video retrieval systems that have to rely both on state-of-the-art ranking models and interactive means of retrieval. We present a general overview of the current version of the VIRET tool, an interactive video retrieval system that successfully participated at several international evaluation campaigns. The system is based on multi-modal search and convenient inspection of results. Based on collected query logs of four users controlling instances of the tool at the Video Browser Showdown 2019, we highlight query modification statistics and a list of successful query formulation strategies. We conclude that the VIRET tool represents a competitive reference interactive system for effective known-item search in one thousand hours of video.
Jakub Lokoc, Gregor Kovalcík, Tomás Soucek, Jaroslav Moravec, Premysl Cech
ICMR3
2019 A Framework for Effective Known-item Search in Video
abstract
Searching for one particular scene in a large video collection (known-item search) represents a challenging task for video retrieval systems. According to the recent results reached at evaluation campaigns, even respected approaches based on machine learning do not help to solve the task easily in many cases. Hence, in addition to effective automatic multimedia annotation and embedding, interactive search is recommended as well. This paper presents a comprehensive description of an interactive video retrieval framework VIRET that successfully participated at several recent evaluation campaigns. Utilized video analysis, feature extraction and retrieval models are detailed as well as several experiments evaluating effectiveness of selected system components. The results of the prototype at the Video Browser Showdown 2019 are highlighted in connection with an analysis of collected query logs. We conclude that the framework comprise a set of effective and efficient models for most of the evaluated known-item search tasks in 1000 hours of video and could serve as a baseline reference approach. The analysis also reveals that the result presentation interface needs improvements for better performance of future VIRET prototypes.
Jakub Lokoc, Gregor Kovalcík, Tomás Soucek, Jaroslav Moravec, Premysl Cech
ACM Multimedia3
2019 VIRET Tool Meets NasNet
Jakub Lokoc, Gregor Kovalcík, Tomás Soucek, Jaroslav Moravec, Jan Bodnár, Premysl Cech
MMM (2)3
2018 Revisiting SIRET Video Retrieval Tool
Jakub Lokoc, Gregor Kovalcík, Tomás Soucek
MMM (2)3