EDBT 2026 Demo / reviewers in the wild / expert
Jan Zahálka
dblp:54/9476
· DBLP profile ↗
22ranked-venue papers
7as first author
8since 2021 · last 2022
0000-0002-6743-3607ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 6 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Exquisitor at the Video Browser Showdown 2022
Omar Shahbaz Khan, Ujjwal Sharma 0001, Björn Þór Jónsson 0001, Dennis C. Koelma, Stevan Rudinac, Marcel Worring, Jan Zahálka |
MMM (2) | 7 |
| 2022 | XQM: Search-Oriented vs. Classifier-Oriented Relevance Feedback on Mobile Phones
Kim I. Schild, Alexandra M. Bagi, Magnus Holm Mamsen, Omar Shahbaz Khan, Jan Zahálka, Björn Þór Jónsson 0001 |
MMM (2) | 5 |
| 2021 | Impact of Interaction Strategies on User Relevance FeedbackabstractUser Relevance Feedback (URF) is a class of interactive learning methods that rely on the interaction between a human user and a system to analyze a media collection. To improve URF system evaluation and design better systems, it is important to understand the impact that different interaction strategies can have. Based on the literature and observations from real user sessions from the Lifelog Search Challenge and Video Browser Showdown, we analyze interaction strategies related to (a) labeling positive and negative examples, and (b) applying filters based on users' domain knowledge. Experiments show that there is no single optimal labeling strategy, as the best strategy depends on both the collection and the task. In particular, our results refute the common assumption that providing more training examples is always beneficial: strategies with a smaller number of prototypical examples lead to better results in some cases. We further observe that while expert filtering is unsurprisingly beneficial, aggressive filtering, especially by novice users, can hinder the completion of tasks. Finally, we observe that combining URF with filters leads to better results than using filters alone. Omar Shahbaz Khan, Björn Þór Jónsson 0001, Jan Zahálka, Stevan Rudinac, Marcel Worring |
ICMR | 3 |
| 2021 | Reproducibility Companion Paper: Kalman Filter-Based Head Motion Prediction for Cloud-Based Mixed RealityabstractIn our MM'20 paper,, we presented a Kalman filter-based approach for prediction of head motion in 6DoF. The proposed approach was employed in our cloud-based volumetric video streaming system to reduce the interaction latency experienced by the user. In this companion paper, we present the dataset collected for our experiments and our simulation framework that reproduces the obtained experimental results. Our implementation is freely available on Github to facilitate further research. Serhan Gul, Sebastian Bosse, Dimitri Podborski, Thomas Schierl, Cornelius Hellge, Marc A. Kastner 0001, Jan Zahálka |
ACM Multimedia | 7 |
| 2021 | Reproducibility Companion Paper: Describing Subjective Experiment Consistency by p-Value P-P PlotabstractIn this paper we reproduce experimental results presented in our earlier work titled "Describing Subjective Experiment Consistency by p-Value P-P Plot" that was presented in the course of the 28th ACM International Conference on Multimedia. The paper aims at verifying the soundness of our prior results and helping others understand our software framework. We present artifacts that help reproduce tables, figures and all the data derived from raw subjective responses that were included in our earlier work. Using the artifacts we show that our results are reproducible. We invite everyone to use our software framework for subjective responses analyses going beyond reproducibility efforts. Jakub Nawala, Lucjan Janowski, Bogdan Cmiel, Krzysztof Rusek, Marc A. Kastner 0001, Jan Zahálka |
ACM Multimedia | 6 |
| 2021 | XQM: Interactive Learning on Mobile Phones
Alexandra M. Bagi, Kim I. Schild, Omar Shahbaz Khan, Jan Zahálka, Björn Þór Jónsson 0001 |
MMM (2) | 4 |
| 2021 | Exquisitor at the Video Browser Showdown 2021: Relationships Between Semantic Classifiers
Omar Shahbaz Khan, Björn Þór Jónsson 0001, Mathias Dybkjær Larsen, Liam Alex Sonto Poulsen, Dennis C. Koelma, Stevan Rudinac, Marcel Worring, Jan Zahálka |
MMM (2) | 8 |
| 2021 | II-20: Intelligent and pragmatic analytic categorization of image collectionsabstractIn this paper, we introduce 11-20 (Image Insight 2020), a multimedia analytics approach for analytic categorization of image collections. Advanced visualizations for image collections exist, but they need tight integration with a machine model to support the task of analytic categorization. Directly employing computer vision and interactive learning techniques gravitates towards search. Analytic categorization, however, is not machine classification (the difference between the two is called the pragmatic gap): a human adds/redefines/deletes categories of relevance on the fly to build insight, whereas the machine classifier is rigid and non-adaptive. Analytic categorization that truly brings the user to insight requires a flexible machine model that allows dynamic sliding on the exploration-search axis, as well as semantic interactions: a human thinks about image data mostly in semantic terms. 11-20 brings three major contributions to multimedia analytics on image collections and towards closing the pragmatic gap. Firstly, a new machine model that closely follows the user's interactions and dynamically models her categories of relevance. II-20's machine model, in addition to matching and exceeding the state of the art's ability to produce relevant suggestions, allows the user to dynamically slide on the exploration-search axis without any additional input from her side. Secondly, the dynamic, 1-image-at-a-time Tetris metaphor that synergizes with the model. It allows a well-trained model to analyze the collection by itself with minimal interaction from the user and complements the classic grid metaphor. Thirdly, the fast-forward interaction, allowing the user to harness the model to quickly expand ("fast-forward") the categories of relevance, expands the multimedia analytics semantic interaction dictionary. Automated experiments show that II-20's machine model outperforms the existing state of the art and also demonstrate the Tetris metaphor's analytic quality. User studies further confirm that II-20 is an intuitive, efficient, and effective multimedia analytics tool. Jan Zahálka, Marcel Worring, Jarke J. van Wijk |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2020 | Interactive Learning for Multimedia at Large
Omar Shahbaz Khan, Björn Þór Jónsson 0001, Stevan Rudinac, Jan Zahálka, Hanna Ragnarsdóttir, Þórhildur Þorleiksdóttir, Gylfi Þór Guðmundsson, Laurent Amsaleg, Marcel Worring |
ECIR (1) | 4 |
| 2020 | Reproducibility Companion Paper: Selective Deep Convolutional Features for Image RetrievalabstractIn this companion paper, firstly, we briefly summarize the contributions of our main manuscript: Selective Deep Convolutional Features for Image Retrieval, published in ACM MultiMedia 2017. In addition, we provide detail instructions together with pre-configured MATLAB scripts which allow experiments to be executed and to reproduce the results reported in our main manuscript effortlessly. The source code is available at https://github.com/hnanhtuan/selectiveConvFeatures_ACMMM_reproducibility. Tuan Hoang, Thanh-Toan Do, Ngai-Man Cheung, Michael Riegler 0001, Jan Zahálka |
ACM Multimedia | 5 |
| 2020 | Reproducibility Companion Paper: Outfit Compatibility Prediction and Diagnosis with Multi-Layered Comparison NetworkabstractThis companion paper supports the experimental replication of paper "Outfit Compatibility Prediction and Diagnosis with Multi-Layered Comparison Network", which is presented at ACM Multimedia 2019. We provide the software package for replicating the implementation of Multi-Layered Comparison Network (MCN), as well as the Polyvore-T dataset and baseline methods compared in the original paper. This paper contains the guides to reproduce the experiment results including outfit compatibility prediction, outfit diagnosis and automatic outfit revision. Xin Wang 0131, Bo Wu 0018, Yueqi Zhong, Wei Hu 0003, Jan Zahálka |
ACM Multimedia | 5 |
| 2020 | Exquisitor at the Video Browser Showdown 2020
Björn Þór Jónsson 0001, Omar Shahbaz Khan, Dennis C. Koelma, Stevan Rudinac, Marcel Worring, Jan Zahálka |
MMM (2) | 6 |
| 2019 | Exquisitor: Breaking the Interaction Barrier for Exploration of 100 Million ImagesabstractIn this demonstration, we present Exquisitor, a media explorer capable of learning user preferences in real-time during interactions with the 99.2 million images of YFCC100M. Exquisitor owes its efficiency to innovations in data representation, compression, and indexing. Exquisitor can complete each interaction round, including learning preferences and presenting the most relevant results, in less than 30 ms using only a single CPU core and modest RAM. In short, Exquisitor can bring large-scale interactive learning to standard desktops and laptops, and even high-end mobile devices. Hanna Ragnarsdóttir, Þórhildur Þorleiksdóttir, Omar Shahbaz Khan, Björn Þór Jónsson 0001, Gylfi Þór Guðmundsson, Jan Zahálka, Stevan Rudinac, Laurent Amsaleg, Marcel Worring |
ACM Multimedia | 6 |
| 2018 | Blackthorn: Large-Scale Interactive Multimodal LearningabstractThis paper presents Blackthorn, an efficient interactive multimodal learning approach facilitating analysis of multimedia collections of up to 100 million items on a single high-end workstation. Blackthorn features efficient data compression, feature selection, and optimizations to the interactive learning process. The Ratio-64 data representation introduced in this paper only costs tens of bytes per item yet preserves most of the visual and textual semantic information with good accuracy. The optimized interactive learning model scores the Ratio-64-compressed data directly, greatly reducing the computational requirements. The experiments compare Blackthorn with two baselines: Conventional relevance feedback, and relevance feedback using product quantization to compress the features. The results show that Blackthorn is up to 77.5× faster than the conventional relevance feedback alternative, while outperforming the baseline with respect to the relevance of results: It vastly outperforms the baseline on recall over time and reaches up to 108% of its precision. Compared to the product quantization variant, Blackthorn is just as fast, while producing more relevant results. On the full YFCC100M dataset, Blackthorn performs one complete interaction round in roughly 1 s while maintaining adequate relevance of results, thus opening multimedia collections comprising up to 100 million items to fully interactive learning-based analysis. Jan Zahálka, Stevan Rudinac, Björn Þór Jónsson 0001, Dennis C. Koelma, Marcel Worring |
IEEE Trans. Multim. | 1 |
| 2017 | Discovering Geographic Regions in the City Using Social Multimedia and Open Data
Stevan Rudinac, Jan Zahálka, Marcel Worring |
MMM (2) | 2 |
| 2016 | Interactive Multimodal Learning on 100 Million ImagesabstractThis paper presents Blackthorn, an efficient interactive multimodal learning approach facilitating analysis of multimedia collections of 100 million items on a single high-end workstation. This is achieved by efficient data compression and optimizations to the interactive learning process. The compressed i-I64 data representation costs tens of bytes per item yet preserves most of the visual and textual semantic information. The optimized interactive learning model scores the i-I64-compressed data directly, greatly reducing the computational requirements. The experiments show that Blackthorn is up to 105x faster than the conventional relevance feedback baseline. Blackthorn is shown to vastly outperform the baseline with respect to recall over time. Blackthorn reaches up to 92% of the precision achieved by the baseline, validating the efficacy of the i-I64 representation. On the YFCC100M dataset, Blackthorn performes one complete interaction round in 0.7 seconds. Blackthorn thus opens multimedia collections comprising 100 million items to learning-based analysis in fully interactive time. Jan Zahálka, Stevan Rudinac, Björn Þór Jónsson 0001, Dennis C. Koelma, Marcel Worring |
ICMR | 1 |
| 2016 | Ten Research Questions for Scalable Multimedia Analytics
Björn Þór Jónsson 0001, Marcel Worring, Jan Zahálka, Stevan Rudinac, Laurent Amsaleg |
MMM (2) | 3 |
| 2016 | Multimedia Pivot Tables for Multimedia Analytics on Image CollectionsabstractWe propose a multimedia analytics solution for getting insight into image collections by extending the powerful analytic capabilities of pivot tables, found in the ubiquitous spreadsheets, to multimedia. We formalize the concept of multimedia pivot tables and give design rules and methods for the multimodal summarization, structuring, and browsing of the collection based on these tables, all optimized to support an analyst in getting structural and conclusive insights. Our proposed solution provides truly interactive analytics on the visual content of image collections through concept detection results, as well as tags, geolocation, time, and other metadata. We have performed user experiments with novice users on a dataset from Flickr to improve the initial design and with expert users in marketing and multimedia analysis on two domain-specific datasets collected from Instagram. The results show that analysts are indeed capable of deriving structural and conclusive insights using the proposed multimedia analytics solution. On our website, videos of the system in action are available. Marcel Worring, Dennis C. Koelma, Jan Zahálka |
IEEE Trans. Multim. | 3 |
| 2015 | Analytic Quality: Evaluation of Performance and Insight in Multimedia Collection AnalysisabstractIn this paper, we present analytic quality (AQ), a novel paradigm for the design and evaluation of multimedia analysis methods. AQ complements the existing evaluation methods based on either machine-driven benchmarks or user studies. AQ includes the notion of user insight gain and the time needed to acquire it, both critical aspects of large-scale multimedia collections analysis. To incorporate insight, AQ introduces a novel user model. In this model, each simulated user, or artificial actor, builds its insight over time, at any time operating with multiple categories of relevance. The methods are evaluated in timed sessions. The artificial actors interact with each method and steer the course by indicating relevant items throughout the session. AQ measures not only precision and recall, but also throughput, diversity of the results, and the accuracy of estimating the percentage of relevant items in the collection. AQ is shown to provide a wide picture of analytic capabilities of the evaluated methods and enumerate how their strengths differ for different purposes. The AQ time plots provide design suggestions for improving the evaluated methods. AQ is demonstrated to be more insightful than the classic benchmark evaluation paradigm both in terms of method comparison and suggestions for further design. Jan Zahálka, Stevan Rudinac, Marcel Worring |
ACM Multimedia | 1 |
| 2015 | Interactive Multimodal Learning for Venue RecommendationabstractIn this paper, we propose City Melange, an interactive and multimodal content-based venue explorer. Our framework matches the interacting user to the users of social media platforms exhibiting similar taste. The data collection integrates location-based social networks such as Foursquare with general multimedia sharing platforms such as Flickr or Picasa. In City Melange, the user interacts with a set of images and thus implicitly with the underlying semantics. The semantic information is captured through convolutional deep net features in the visual domain and latent topics extracted using Latent Dirichlet allocation in the text domain. These are further clustered to provide representative user and venue topics. A linear SVM model learns the interacting user's preferences and determines similar users. The experiments show that our content-based approach outperforms the user-activity-based and popular vote baselines even from the early phases of interaction, while also being able to recommend mainstream venues to mainstream users and off-the-beaten-track venues to afficionados. City Melange is shown to be a well-performing venue exploration approach. Jan Zahálka, Stevan Rudinac, Marcel Worring |
IEEE Trans. Multim. | 1 |
| 2014 | New Yorker Melange: Interactive Brew of Personalized Venue RecommendationsabstractIn this paper we propose New Yorker Melange, an interactive city explorer, which navigates New York venues through the eyes of New Yorkers having a similar taste to the interacting user. To gain insight into New Yorkers' preferences and properties of the venues, a dataset of more than a million venue images and associated annotations has been collected from Foursquare, Picasa, and Flickr. As visual and text features, we use semantic concepts extracted by a convolutional deep net and latent Dirichlet allocation topics. To identify different aspects of the venues and topics of interest to the users, we further cluster images associated with them. New Yorker Melange uses an interactive map interface and learns the interacting user's taste using linear SVM. The SVM model is used to navigate the interacting user's exploration further towards similar users. Experimental evaluation demonstrates that our proposed approach is effective in producing relevant results and that both visual and text modalities contribute to the overall system performance. Jan Zahálka, Stevan Rudinac, Marcel Worring |
ACM Multimedia | 1 |
| 2011 | An experimental test of Occam's razor in classification
Jan Zahálka, Filip Zelezný |
Mach. Learn. | 1 |