EDBT 2026 Demo / reviewers in the wild / expert
Michael Gygli
dblp:135/4875
· DBLP profile ↗
20ranked-venue papers
10as first author
3since 2021 · last 2024
0000-0001-5241-481XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 8 first-author · 2 since 2021Artificial intelligence and machine learning · 15 · 9 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
13 papers |
Transfer learning and domain adaptation · 20% Segmentation and scene understanding · 18% Representation and self-supervised learning · 12% | |
| Computer graphics and multimedia
7 papers |
Multimedia analysis and retrieval · 79% Audio and music processing · 21% | |
| Human-computer interaction and pervasive computing
2 papers |
Haptics and multimodal interaction · 54% Interaction techniques and input · 46% |
Topics — the 30 heaviest of 43, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Multimedia analysis and retrieval
video summarization |
0.8 | 3 | 2018 | PHD-GIFs: Personalized Highlight Detection for Automatic GIF Creation · ACM Multimedia 2018 Query-adaptive Video Summarization via Quality-aware Relevance Estimation · ACM Multimedia 2017 Video summarization by learning submodular mixtures of objectives · CVPR 2015 |
Multimedia analysis and retrieval › video summarization
video highlight detection |
0.6 | 2 | 2018 | PHD-GIFs: Personalized Highlight Detection for Automatic GIF Creation · ACM Multimedia 2018 Video2GIF: Automatic Generation of Animated GIFs from Video · CVPR 2016 |
Machine learning › Transfer learning and domain adaptation
cross-domain learning |
0.6 | 1 | 2022 | Factors of Influence for Transfer Learning Across Diverse Appearance Domains and Task Types · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Machine learning › Transfer learning and domain adaptation
cross-task transfer |
0.6 | 1 | 2022 | Factors of Influence for Transfer Learning Across Diverse Appearance Domains and Task Types · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Machine learning › Representation and self-supervised learning
compatible representations |
0.5 | 1 | 2021 | Towards Reusable Network Components by Learning Compatible Representations · AAAI 2021 |
Machine learning › Transfer learning and domain adaptation › model adaptation › online adaptation
continuous adaptation |
0.4 | 1 | 2020 | Continuous Adaptation for Interactive Object Segmentation by Learning from Corrections · ECCV (16) 2020 |
Computer vision › Segmentation and scene understanding › interactive segmentation
interactive object segmentation |
0.4 | 1 | 2020 | Continuous Adaptation for Interactive Object Segmentation by Learning from Corrections · ECCV (16) 2020 |
Computer vision › Segmentation and scene understanding
interactive segmentation |
0.4 | 1 | 2020 | Continuous Adaptation for Interactive Object Segmentation by Learning from Corrections · ECCV (16) 2020 |
Robotics › Robot manipulation › learning from demonstration
learning from corrections |
0.4 | 1 | 2020 | Continuous Adaptation for Interactive Object Segmentation by Learning from Corrections · ECCV (16) 2020 |
Haptics and multimodal interaction
multimodal interaction |
0.4 | 1 | 2020 | Efficient Object Annotation via Speaking and Pointing · Int. J. Comput. Vis. 2020 |
Interaction techniques and input
voice interaction |
0.4 | 1 | 2019 | Fast Object Class Labelling via Speech · CVPR 2019 |
Computer vision › Video understanding and tracking
action recognition |
0.3 | 1 | 2018 | AENet: Learning Deep Audio Features for Video Analysis · IEEE Trans. Multim. 2018 |
Audio and music processing
sound event detection |
0.3 | 1 | 2018 | AENet: Learning Deep Audio Features for Video Analysis · IEEE Trans. Multim. 2018 |
Audio and music processing › sound event detection
sound event recognition |
0.3 | 1 | 2018 | AENet: Learning Deep Audio Features for Video Analysis · IEEE Trans. Multim. 2018 |
Natural language and speech › Information extraction and text analysis
dataset construction |
0.3 | 1 | 2017 | PathTrack: Fast Trajectory Annotation with Path Supervision · ICCV 2017 |
Computer vision › Segmentation and scene understanding
image segmentation |
0.3 | 1 | 2017 | Deep Value Networks Learn to Evaluate and Iteratively Refine Structured Outputs · ICML 2017 |
Computer vision › Video understanding and tracking
multi-object tracking |
0.3 | 1 | 2017 | PathTrack: Fast Trajectory Annotation with Path Supervision · ICCV 2017 |
Computer vision › Segmentation and scene understanding › semantic segmentation
segmentation refinement |
0.3 | 1 | 2017 | Deep Value Networks Learn to Evaluate and Iteratively Refine Structured Outputs · ICML 2017 |
Machine learning › Probabilistic and Bayesian machine learning
structured prediction |
0.3 | 1 | 2017 | Deep Value Networks Learn to Evaluate and Iteratively Refine Structured Outputs · ICML 2017 |
Multimedia analysis and retrieval
cross-modal retrieval |
0.3 | 1 | 2017 | Query-adaptive Video Summarization via Quality-aware Relevance Estimation · ACM Multimedia 2017 |
Multimedia analysis and retrieval › video summarization
query-focused video summarization |
0.3 | 1 | 2017 | Query-adaptive Video Summarization via Quality-aware Relevance Estimation · ACM Multimedia 2017 |
Multimedia analysis and retrieval › multimedia analysis › cross-media analysis › vision-language understanding
visual-semantic embedding |
0.3 | 1 | 2017 | Query-adaptive Video Summarization via Quality-aware Relevance Estimation · ACM Multimedia 2017 |
Machine learning › Deep learning architectures and training
deep ranking |
0.2 | 1 | 2016 | Video2GIF: Automatic Generation of Animated GIFs from Video · CVPR 2016 |
Machine learning › Trustworthy machine learning
interpretability |
0.2 | 1 | 2016 | Predicting When Saliency Maps are Accurate and Eye Fixations Consistent · CVPR 2016 |
Machine learning › Learning theory › ranking
learning to rank |
0.2 | 1 | 2016 | Video2GIF: Automatic Generation of Animated GIFs from Video · CVPR 2016 |
Computer vision › Image recognition and object detection
saliency prediction |
0.2 | 1 | 2016 | Predicting When Saliency Maps are Accurate and Eye Fixations Consistent · CVPR 2016 |
Web and social media mining › social media analysis
social media content analysis |
0.2 | 1 | 2016 | Analyzing and Predicting GIF Interestingness · ACM Multimedia 2016 |
Computer vision › Image recognition and object detection
object labeling |
0.2 | 2 | 2020 | Efficient Object Annotation via Speaking and Pointing · Int. J. Comput. Vis. 2020 Fast Object Class Labelling via Speech · CVPR 2019 |
Machine learning › Optimization for machine learning › combinatorial optimization
submodular optimization |
0.2 | 1 | 2015 | Video summarization by learning submodular mixtures of objectives · CVPR 2015 |
Computer vision › Video understanding and tracking
video summarization |
0.2 | 1 | 2014 | Creating Summaries from User Videos · ECCV (7) 2014 |
Methods — techniques the papers use, named apart from their topics
speech recognition · 1.6pointing gesture recognition · 0.9feature engineering · 0.7global ranking model · 0.7data augmentation · 0.7convolutional neural network · 0.7pre-training and fine-tuning · 0.6empirical benchmarking · 0.6predictive modeling · 0.5fine-tuning · 0.5feature extraction · 0.5interactive learning · 0.4correction-based adaptation · 0.4transfer learning · 0.3neural network · 0.3maximal marginal relevance · 0.3ranknet · 0.2adaptive huber loss · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | CycleCL: Self-supervised Learning for Periodic VideosabstractAnalyzing periodic video sequences is a key topic in applications such as automatic production systems, remote sensing, medical applications, or physical training. An example is counting repetitions of a physical exercise. Due to the distinct characteristics of periodic data, self-supervised methods designed for standard image datasets do not capture changes relevant to the progression of the cycle and fail to ignore unrelated noise. They thus do not work well on periodic data.In this paper, we propose CycleCL, a self-supervised learning method specifically designed to work with periodic data. We start from the insight that a good visual representation for periodic data should be sensitive to the phase of a cycle, but be invariant to the exact repetition, i.e. it should generate identical representations for a specific phase throughout all repetitions. We exploit the repetitions in videos to design a novel contrastive learning method based on a triplet loss that optimizes for these desired properties. Our method uses pre-trained features to sample pairs of frames from approximately the same phase and negative pairs of frames from different phases. Then, we iterate between optimizing a feature encoder and resampling triplets, until convergence.By optimizing a model this way, we are able to learn features that have the mentioned desired properties. We evaluate CycleCL on an industrial and multiple human actions datasets, where it significantly outperforms previous video-based self-supervised learning methods on all tasks. Matteo Destro, Michael Gygli |
WACV | 2 |
| 2022 | Factors of Influence for Transfer Learning Across Diverse Appearance Domains and Task TypesabstractTransfer learning enables to re-use knowledge learned on a source task to help learning a target task. A simple form of transfer learning is common in current state-of-the-art computer vision models, i.e., pre-training a model for image classification on the ILSVRC dataset, and then fine-tune on any target task. However, previous systematic studies of transfer learning have been limited and the circumstances in which it is expected to work are not fully understood. In this paper we carry out an extensive experimental exploration of transfer learning across vastly different image domains (consumer photos, autonomous driving, aerial imagery, underwater, indoor scenes, synthetic, close-ups) and task types (semantic segmentation, object detection, depth estimation, keypoint detection). Importantly, these are all complex, structured output tasks types relevant to modern computer vision applications. In total we carry out over 2000 transfer learning experiments, including many where the source and target come from different image domains, task types, or both. We systematically analyze these experiments to understand the impact of image domain, task type, and dataset size on transfer learning performance. Our study leads to several insights and concrete recommendations: (1) for most tasks there exists a source which significantly outperforms ILSVRC'12 pre-training; (2) the image domain is the most important factor for achieving positive transfer; (3) the source dataset should include the image domain of the target dataset to achieve best results; (4) at the same time, we observe only small negative effects when the image domain of the source task is much broader than that of the target; (5) transfer across task types can be beneficial, but its success is heavily dependent on both the source and target task types. Thomas Mensink, Jasper R. R. Uijlings, Alina Kuznetsova, Michael Gygli, Vittorio Ferrari |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Towards Reusable Network Components by Learning Compatible RepresentationsabstractThis paper proposes to make a first step towards compatible and hence reusable network components. Rather than training networks for different tasks independently, we adapt the training process to produce network components that are compatible across tasks. In particular, we split a network into two components, a features extractor and a target task head, and propose various approaches to accomplish compatibility between them. We systematically analyse these approaches on the task of image classification on standard datasets. We demonstrate that we can produce components which are directly compatible without any fine-tuning or compromising accuracy on the original tasks. Afterwards, we demonstrate the use of compatible components on three applications: Unsupervised domain adaptation, transferring classifiers across feature extractors with different architectures, and increasing the computational efficiency of transfer learning. Michael Gygli, Jasper R. R. Uijlings, Vittorio Ferrari |
AAAI | 1 |
| 2020 | Continuous Adaptation for Interactive Object Segmentation by Learning from Corrections
Theodora Kontogianni, Michael Gygli, Jasper R. R. Uijlings, Vittorio Ferrari |
ECCV (16) | 2 |
| 2020 | Efficient Object Annotation via Speaking and Pointing
Michael Gygli, Vittorio Ferrari |
Int. J. Comput. Vis. | 1 |
| 2019 | Fast Object Class Labelling via SpeechabstractObject class labelling is the task of annotating images with labels on the presence or absence of objects from a given class vocabulary. Simply asking one yes-no question per class, however, has a cost that is linear in the vocabulary size and is thus inefficient for large vocabularies. Modern approaches rely on a hierarchical organization of the vocabulary to reduce annotation time, but remain expensive (several minutes per image for the 200 classes in ILSVRC). Instead, we propose a new interface where classes are annotated via speech. Speaking is fast and allows for direct access to the class name, without searching through a list or hierarchy. As additional advantages, annotators can simultaneously speak and scan the image for objects, the interface can be kept extremely simple, and using it requires less mouse movement. As annotators using our interface should only say words from a given class vocabulary, we propose a dedicated task to train them to do so. Through experiments on COCO and ILSVRC, we show our method yields high-quality annotations at 2.3x -14.9x less annotation time than existing methods. Michael Gygli, Vittorio Ferrari |
CVPR | 1 |
| 2018 | Ridiculously Fast Shot Boundary Detection with Fully Convolutional Neural NetworksabstractShot boundary detection (SBD) is an important component of many video analysis tasks, such as action recognition' video indexing, summarization and editing. Previous work typically used a combination of low-level features like color histograms, in conjunction with simple models such as SVMs to predict shot changes. Instead, we propose to learn shot detection end-to-end, from pixels to final shot boundaries. For training such a model, we rely on our insight that all shot boundaries are generated. Thus, we create a dataset with one million frames and automatically generated transitions such as cuts, dissolves and fades. In order to efficiently analyze hours of videos, we propose a Convolutional Neural Network (CNN) which is fully convolutional in time, thus allowing to use a large temporal context without the need to repeatedly processing frames. With this architecture our method obtains state-of-the-art results on the RAI dataset, while running at an unprecedented speed of more than 120x real-time. Michael Gygli |
CBMI | 1 |
| 2018 | PHD-GIFs: Personalized Highlight Detection for Automatic GIF CreationabstractHighlight detection models are typically trained to identify cues that make visual content appealing or interesting for the general public, with the objective of reducing a video to such moments. However, this "interestingness" of a video segment or image is subjective. Thus, such highlight models provide results of limited relevance for the individual user. On the other hand, training one model per user is inefficient and requires large amounts of personal information which is typically not available. To overcome these limitations, we present a global ranking model which can condition on a particular user's interests. Rather than training one model per user, our model is personalized via its inputs, which allows it to effectively adapt its predictions, given only a few user-specific examples. To train this model, we create a large-scale dataset of users and the GIFs they created, giving us an accurate indication of their interests. Our experiments show that using the user history substantially improves the prediction accuracy. On a test set of 850 videos, our model improves the recall by 8% with respect to generic highlight detectors. Furthermore, our method proves more precise than the user-agnostic baselines even with only one single person-specific example. Ana Garcia del Molino, Michael Gygli |
ACM Multimedia | 2 |
| 2018 | AENet: Learning Deep Audio Features for Video AnalysisabstractWe propose a new deep network for audio event recognition, called AENet. In contrast to speech, sounds coming from audio events may be produced by a wide variety of sources. Furthermore, distinguishing them often requires analyzing an extended time period due to the lack of clear subword units that are present in speech. In order to incorporate this long-time frequency structure of audio events, we introduce a convolutional neural network (CNN) operating on a large temporal input. In contrast to previous works, this allows us to train an audio event detection system end to end. The combination of our network architecture and a novel data augmentation outperforms previous methods for audio event detection by 16%. Furthermore, we perform transfer learning and show that our model learned generic audio features, similar to the way CNNs learn generic features on vision tasks. In video analysis, combining visual features and traditional audio features, such as mel frequency cepstral coefficients, typically only leads to marginal improvements. Instead, combining visual features with our AENet features, which can be computed efficiently on a GPU, leads to significant performance improvements on action recognition and video highlight detection. In video highlight detection, our audio features improve the performance by more than 8% over visual features alone. Naoya Takahashi, Michael Gygli, Luc Van Gool |
IEEE Trans. Multim. | 2 |
| 2017 | PathTrack: Fast Trajectory Annotation with Path SupervisionabstractProgress in Multiple Object Tracking (MOT) has been historically limited by the size of the available datasets. We present an efficient framework to annotate trajectories and use it to produce a MOT dataset of unprecedented size. In our novel path supervision the annotator loosely follows the object with the cursor while watching the video, providing a path annotation for each object in the sequence. Our approach is able to turn such weak annotations into dense box trajectories. Our experiments on existing datasets prove that our framework produces more accurate annotations than the state of the art, in a fraction of the time. We further validate our approach by crowdsourcing the PathTrack dataset, with more than 15,000 person trajectories in 720 sequences. Tracking approaches can benefit training on such large-scale datasets, as did object recognition. We prove this by re-training an off-the-shelf person matching network, originally trained on the MOT15 dataset, almost halving the misclassification rate. Additionally, training on our data consistently improves tracking results, both on our dataset and on MOT15. On the latter, we improve the top-performing tracker (NOMT) dropping the number of ID Switches by 18% and fragments by 5%. Santiago Manen, Michael Gygli, Dengxin Dai, Luc Van Gool |
ICCV | 2 |
| 2017 | Deep Value Networks Learn to Evaluate and Iteratively Refine Structured OutputsabstractWe approach structured output prediction by optimizing a deep value network (DVN) to precisely estimate the task loss on different output configurations for a given input. Once the model is trained, we perform inference by gradient descent on the continuous relaxations of the output variables to find outputs with promising scores from the value network. When applied to image segmentation, the value network takes an image and a segmentation mask as inputs and predicts a scalar estimating the intersection over union between the input and ground truth masks. For multi-label classification, the DVN’s objective is to correctly predict the F1 score for any potential label configuration. The DVN framework achieves the state-of-the-art results on multi-label prediction and image segmentation benchmarks. Michael Gygli, Mohammad Norouzi 0002, Anelia Angelova |
ICML | 1 |
| 2017 | Query-adaptive Video Summarization via Quality-aware Relevance EstimationabstractAlthough the problem of automatic video summarization has recently received a lot of attention, the problem of creating a video summary that also highlights elements relevant to a search query has been less studied. We address this problem by posing query-relevant summarization as a video frame subset selection problem, which lets us optimise for summaries which are simultaneously diverse, representative of the entire video, and relevant to a text query. We quantify relevance by measuring the distance between frames and queries in a common textual-visual semantic embedding space induced by a neural network. In addition, we extend the model to capture query-independent properties, such as frame quality. We compare our method against previous state of the art on textual-visual embeddings for thumbnail selection and show that our model outperforms them on relevance prediction. Furthermore, we introduce a new dataset, annotated with diversity and query-specific relevance labels. On this dataset, we train and test our complete model for video summarization and show that it outperforms standard baselines such as Maximal Marginal Relevance. Arun Balajee Vasudevan, Michael Gygli, Anna Volokitin, Luc Van Gool |
ACM Multimedia | 2 |
| 2016 | Video2GIF: Automatic Generation of Animated GIFs from VideoabstractWe introduce the novel problem of automatically generating animated GIFs from video. GIFs are short looping video with no sound, and a perfect combination between image and video that really capture our attention. GIFs tell a story, express emotion, turn events into humorous moments, and are the new wave of photojournalism. We pose the question: Can we automate the entirely manual and elaborate process of GIF creation by leveraging the plethora of user generated GIF content? We propose a Robust Deep RankNet that, given a video, generates a ranked list of its segments according to their suitability as GIF. We train our model to learn what visual content is often selected for GIFs by using over 100K user generated GIFs and their corresponding video sources. We effectively deal with the noisy web data by proposing a novel adaptive Huber loss in the ranking formulation. We show that our approach is robust to outliers and picks up several patterns that are frequently present in popular animated GIFs. On our new large-scale benchmark dataset, we show the advantage of our approach over several state-of-the-art methods. Michael Gygli, Yale Song, Liangliang Cao |
CVPR | 1 |
| 2016 | Predicting When Saliency Maps are Accurate and Eye Fixations ConsistentabstractMany computational models of visual attention use image features and machine learning techniques to predict eye fixation locations as saliency maps. Recently, the success of Deep Convolutional Neural Networks (DCNNs) for object recognition has opened a new avenue for computational models of visual attention due to the tight link between visual attention and object recognition. In this paper, we show that using features from DCNNs for object recognition we can make predictions that enrich the information provided by saliency models. Namely, we can estimate the reliability of a saliency model from the raw image, which serves as a meta-saliency measure that may be used to select the best saliency algorithm for an image. Analogously, the consistency of the eye fixations among subjects, i.e. the agreement between the eye fixation locations of different subjects, can also be predicted and used by a designer to assess whether subjects reach a consensus about salient image locations. Anna Volokitin, Michael Gygli, Xavier Boix |
CVPR | 2 |
| 2016 | Deep Convolutional Neural Networks and Data Augmentation for Acoustic Event Recognition
Naoya Takahashi, Michael Gygli, Beat Pfister, Luc Van Gool |
INTERSPEECH | 2 |
| 2016 | Analyzing and Predicting GIF InterestingnessabstractAnimated GIFs have regained huge popularity. They are used in instant messaging, online journalism, social media, among others. In this paper, we present an in-depth study on the interestingness of GIFs. We create and annotate a dataset with a set of affective labels, which allows us to investigate the sources of interest. We show that GIFs of pets are considered more interesting that GIFs of people. Furthermore, we study the connection of interest to other features and factors such as popularity. Finally, we build a predictive model and show that it can estimate GIF interestingness with high accuracy. Our model outperforms the existing methods on GIF popularity, as well as a model based on still image interestingness, by a large margin. We envision that the insights and method developed can be used for automatic recognition and generation of interesting GIFs. Michael Gygli, Mohammad Soleymani 0001 |
ACM Multimedia | 1 |
| 2015 | Video summarization by learning submodular mixtures of objectivesabstractWe present a novel method for summarizing raw, casually captured videos. The objective is to create a short summary that still conveys the story. It should thus be both, interesting and representative for the input video. Previous methods often used simplified assumptions and only optimized for one of these goals. Alternatively, they used handdefined objectives that were optimized sequentially by making consecutive hard decisions. This limits their use to a particular setting. Instead, we introduce a new method that (i) uses a supervised approach in order to learn the importance of global characteristics of a summary and (ii) jointly optimizes for multiple objectives and thus creates summaries that posses multiple properties of a good summary. Experiments on two challenging and very diverse datasets demonstrate the effectiveness of our method, where we outperform or match current state-of-the-art. Michael Gygli, Helmut Grabner, Luc Van Gool |
CVPR | 1 |
| 2014 | Creating Summaries from User Videos
Michael Gygli, Helmut Grabner, Hayko Riemenschneider, Luc Van Gool |
ECCV (7) | 1 |
| 2013 | Sparse Quantization for Patch DescriptionabstractThe representation of local image patches is crucial for the good performance and efficiency of many vision tasks. Patch descriptors have been designed to generalize towards diverse variations, depending on the application, as well as the desired compromise between accuracy and efficiency. We present a novel formulation of patch description, that serves such issues well. Sparse quantization lies at its heart. This allows for efficient encodings, leading to powerful, novel binary descriptors, yet also to the generalization of existing descriptors like SIFT or BRIEF. We demonstrate the capabilities of our formulation for both key point matching and image classification. Our binary descriptors achieve state-of-the-art results for two key point matching benchmarks, namely those by Brown and Mikolajczyk. For image classification, we propose new descriptors, that perform similar to SIFT on Caltech101 and PASCAL VOC07. Xavier Boix, Michael Gygli, Gemma Roig, Luc Van Gool |
CVPR | 2 |
| 2013 | The Interestingness of ImagesabstractWe investigate human interest in photos. Based on our own and others' psychophysical experiments, we identify various cues for "interestingness", namely aesthetics, unusualness and general preferences. For the ranking of retrieved images, interestingness shows to be more appropriate than cues proposed earlier. Interestingness is correlated with what people believe they will remember. This is opposed to actual memorability, which is uncorrelated to both. We introduce a set of features computationally capturing the three main aspects of visual interestingness and build an interestingness predictor from them. Its performance is shown on three datasets with varying context, reflecting the prior knowledge of the viewers. Michael Gygli, Helmut Grabner, Hayko Riemenschneider, Fabian Nater, Luc Van Gool |
ICCV | 1 |