Michael Gygli

dblp:135/4875 · DBLP profile ↗
← Back
20ranked-venue papers
10as first author
3since 2021 · last 2024
0000-0001-5241-481XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 8 first-author · 2 since 2021Artificial intelligence and machine learning · 15 · 9 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
13 papers
Transfer learning and domain adaptation · 20% Segmentation and scene understanding · 18% Representation and self-supervised learning · 12%
Computer graphics and multimedia
7 papers
Multimedia analysis and retrieval · 79% Audio and music processing · 21%
Human-computer interaction and pervasive computing
2 papers
Haptics and multimodal interaction · 54% Interaction techniques and input · 46%

Topics — the 30 heaviest of 43, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Multimedia analysis and retrieval
video summarization
0.832018
PHD-GIFs: Personalized Highlight Detection for Automatic GIF Creation · ACM Multimedia 2018
Query-adaptive Video Summarization via Quality-aware Relevance Estimation · ACM Multimedia 2017
Video summarization by learning submodular mixtures of objectives · CVPR 2015
Multimedia analysis and retrieval › video summarization
video highlight detection
0.622018
PHD-GIFs: Personalized Highlight Detection for Automatic GIF Creation · ACM Multimedia 2018
Video2GIF: Automatic Generation of Animated GIFs from Video · CVPR 2016
Machine learning › Transfer learning and domain adaptation
cross-domain learning
0.612022
Factors of Influence for Transfer Learning Across Diverse Appearance Domains and Task Types · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Machine learning › Transfer learning and domain adaptation
cross-task transfer
0.612022
Factors of Influence for Transfer Learning Across Diverse Appearance Domains and Task Types · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Machine learning › Representation and self-supervised learning
compatible representations
0.512021
Towards Reusable Network Components by Learning Compatible Representations · AAAI 2021
Machine learning › Transfer learning and domain adaptation › model adaptation › online adaptation
continuous adaptation
0.412020
Continuous Adaptation for Interactive Object Segmentation by Learning from Corrections · ECCV (16) 2020
Computer vision › Segmentation and scene understanding › interactive segmentation
interactive object segmentation
0.412020
Continuous Adaptation for Interactive Object Segmentation by Learning from Corrections · ECCV (16) 2020
Computer vision › Segmentation and scene understanding
interactive segmentation
0.412020
Continuous Adaptation for Interactive Object Segmentation by Learning from Corrections · ECCV (16) 2020
Robotics › Robot manipulation › learning from demonstration
learning from corrections
0.412020
Continuous Adaptation for Interactive Object Segmentation by Learning from Corrections · ECCV (16) 2020
Haptics and multimodal interaction
multimodal interaction
0.412020
Efficient Object Annotation via Speaking and Pointing · Int. J. Comput. Vis. 2020
Interaction techniques and input
voice interaction
0.412019
Fast Object Class Labelling via Speech · CVPR 2019
Computer vision › Video understanding and tracking
action recognition
0.312018
AENet: Learning Deep Audio Features for Video Analysis · IEEE Trans. Multim. 2018
Audio and music processing
sound event detection
0.312018
AENet: Learning Deep Audio Features for Video Analysis · IEEE Trans. Multim. 2018
Audio and music processing › sound event detection
sound event recognition
0.312018
AENet: Learning Deep Audio Features for Video Analysis · IEEE Trans. Multim. 2018
Natural language and speech › Information extraction and text analysis
dataset construction
0.312017
PathTrack: Fast Trajectory Annotation with Path Supervision · ICCV 2017
Computer vision › Segmentation and scene understanding
image segmentation
0.312017
Deep Value Networks Learn to Evaluate and Iteratively Refine Structured Outputs · ICML 2017
Computer vision › Video understanding and tracking
multi-object tracking
0.312017
PathTrack: Fast Trajectory Annotation with Path Supervision · ICCV 2017
Computer vision › Segmentation and scene understanding › semantic segmentation
segmentation refinement
0.312017
Deep Value Networks Learn to Evaluate and Iteratively Refine Structured Outputs · ICML 2017
Machine learning › Probabilistic and Bayesian machine learning
structured prediction
0.312017
Deep Value Networks Learn to Evaluate and Iteratively Refine Structured Outputs · ICML 2017
Multimedia analysis and retrieval
cross-modal retrieval
0.312017
Query-adaptive Video Summarization via Quality-aware Relevance Estimation · ACM Multimedia 2017
Multimedia analysis and retrieval › video summarization
query-focused video summarization
0.312017
Query-adaptive Video Summarization via Quality-aware Relevance Estimation · ACM Multimedia 2017
Multimedia analysis and retrieval › multimedia analysis › cross-media analysis › vision-language understanding
visual-semantic embedding
0.312017
Query-adaptive Video Summarization via Quality-aware Relevance Estimation · ACM Multimedia 2017
Machine learning › Deep learning architectures and training
deep ranking
0.212016
Video2GIF: Automatic Generation of Animated GIFs from Video · CVPR 2016
Machine learning › Trustworthy machine learning
interpretability
0.212016
Predicting When Saliency Maps are Accurate and Eye Fixations Consistent · CVPR 2016
Machine learning › Learning theory › ranking
learning to rank
0.212016
Video2GIF: Automatic Generation of Animated GIFs from Video · CVPR 2016
Computer vision › Image recognition and object detection
saliency prediction
0.212016
Predicting When Saliency Maps are Accurate and Eye Fixations Consistent · CVPR 2016
Web and social media mining › social media analysis
social media content analysis
0.212016
Analyzing and Predicting GIF Interestingness · ACM Multimedia 2016
Computer vision › Image recognition and object detection
object labeling
0.222020
Efficient Object Annotation via Speaking and Pointing · Int. J. Comput. Vis. 2020
Fast Object Class Labelling via Speech · CVPR 2019
Machine learning › Optimization for machine learning › combinatorial optimization
submodular optimization
0.212015
Video summarization by learning submodular mixtures of objectives · CVPR 2015
Computer vision › Video understanding and tracking
video summarization
0.212014
Creating Summaries from User Videos · ECCV (7) 2014

Methods — techniques the papers use, named apart from their topics

speech recognition · 1.6pointing gesture recognition · 0.9feature engineering · 0.7global ranking model · 0.7data augmentation · 0.7convolutional neural network · 0.7pre-training and fine-tuning · 0.6empirical benchmarking · 0.6predictive modeling · 0.5fine-tuning · 0.5feature extraction · 0.5interactive learning · 0.4correction-based adaptation · 0.4transfer learning · 0.3neural network · 0.3maximal marginal relevance · 0.3ranknet · 0.2adaptive huber loss · 0.2
YearPublicationVenuePosition
2024 CycleCL: Self-supervised Learning for Periodic Videos
abstract
Analyzing periodic video sequences is a key topic in applications such as automatic production systems, remote sensing, medical applications, or physical training. An example is counting repetitions of a physical exercise. Due to the distinct characteristics of periodic data, self-supervised methods designed for standard image datasets do not capture changes relevant to the progression of the cycle and fail to ignore unrelated noise. They thus do not work well on periodic data.In this paper, we propose CycleCL, a self-supervised learning method specifically designed to work with periodic data. We start from the insight that a good visual representation for periodic data should be sensitive to the phase of a cycle, but be invariant to the exact repetition, i.e. it should generate identical representations for a specific phase throughout all repetitions. We exploit the repetitions in videos to design a novel contrastive learning method based on a triplet loss that optimizes for these desired properties. Our method uses pre-trained features to sample pairs of frames from approximately the same phase and negative pairs of frames from different phases. Then, we iterate between optimizing a feature encoder and resampling triplets, until convergence.By optimizing a model this way, we are able to learn features that have the mentioned desired properties. We evaluate CycleCL on an industrial and multiple human actions datasets, where it significantly outperforms previous video-based self-supervised learning methods on all tasks.
Matteo Destro, Michael Gygli
WACV2
2022 Factors of Influence for Transfer Learning Across Diverse Appearance Domains and Task Types
abstract
Transfer learning enables to re-use knowledge learned on a source task to help learning a target task. A simple form of transfer learning is common in current state-of-the-art computer vision models, i.e., pre-training a model for image classification on the ILSVRC dataset, and then fine-tune on any target task. However, previous systematic studies of transfer learning have been limited and the circumstances in which it is expected to work are not fully understood. In this paper we carry out an extensive experimental exploration of transfer learning across vastly different image domains (consumer photos, autonomous driving, aerial imagery, underwater, indoor scenes, synthetic, close-ups) and task types (semantic segmentation, object detection, depth estimation, keypoint detection). Importantly, these are all complex, structured output tasks types relevant to modern computer vision applications. In total we carry out over 2000 transfer learning experiments, including many where the source and target come from different image domains, task types, or both. We systematically analyze these experiments to understand the impact of image domain, task type, and dataset size on transfer learning performance. Our study leads to several insights and concrete recommendations: (1) for most tasks there exists a source which significantly outperforms ILSVRC'12 pre-training; (2) the image domain is the most important factor for achieving positive transfer; (3) the source dataset should include the image domain of the target dataset to achieve best results; (4) at the same time, we observe only small negative effects when the image domain of the source task is much broader than that of the target; (5) transfer across task types can be beneficial, but its success is heavily dependent on both the source and target task types.
Thomas Mensink, Jasper R. R. Uijlings, Alina Kuznetsova, Michael Gygli, Vittorio Ferrari
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 Towards Reusable Network Components by Learning Compatible Representations
abstract
This paper proposes to make a first step towards compatible and hence reusable network components. Rather than training networks for different tasks independently, we adapt the training process to produce network components that are compatible across tasks. In particular, we split a network into two components, a features extractor and a target task head, and propose various approaches to accomplish compatibility between them. We systematically analyse these approaches on the task of image classification on standard datasets. We demonstrate that we can produce components which are directly compatible without any fine-tuning or compromising accuracy on the original tasks. Afterwards, we demonstrate the use of compatible components on three applications: Unsupervised domain adaptation, transferring classifiers across feature extractors with different architectures, and increasing the computational efficiency of transfer learning.
Michael Gygli, Jasper R. R. Uijlings, Vittorio Ferrari
AAAI1
2020 Continuous Adaptation for Interactive Object Segmentation by Learning from Corrections
Theodora Kontogianni, Michael Gygli, Jasper R. R. Uijlings, Vittorio Ferrari
ECCV (16)2
2020 Efficient Object Annotation via Speaking and Pointing
Michael Gygli, Vittorio Ferrari
Int. J. Comput. Vis.1
2019 Fast Object Class Labelling via Speech
abstract
Object class labelling is the task of annotating images with labels on the presence or absence of objects from a given class vocabulary. Simply asking one yes-no question per class, however, has a cost that is linear in the vocabulary size and is thus inefficient for large vocabularies. Modern approaches rely on a hierarchical organization of the vocabulary to reduce annotation time, but remain expensive (several minutes per image for the 200 classes in ILSVRC). Instead, we propose a new interface where classes are annotated via speech. Speaking is fast and allows for direct access to the class name, without searching through a list or hierarchy. As additional advantages, annotators can simultaneously speak and scan the image for objects, the interface can be kept extremely simple, and using it requires less mouse movement. As annotators using our interface should only say words from a given class vocabulary, we propose a dedicated task to train them to do so. Through experiments on COCO and ILSVRC, we show our method yields high-quality annotations at 2.3x -14.9x less annotation time than existing methods.
Michael Gygli, Vittorio Ferrari
CVPR1
2018 Ridiculously Fast Shot Boundary Detection with Fully Convolutional Neural Networks
abstract
Shot boundary detection (SBD) is an important component of many video analysis tasks, such as action recognition' video indexing, summarization and editing. Previous work typically used a combination of low-level features like color histograms, in conjunction with simple models such as SVMs to predict shot changes. Instead, we propose to learn shot detection end-to-end, from pixels to final shot boundaries. For training such a model, we rely on our insight that all shot boundaries are generated. Thus, we create a dataset with one million frames and automatically generated transitions such as cuts, dissolves and fades. In order to efficiently analyze hours of videos, we propose a Convolutional Neural Network (CNN) which is fully convolutional in time, thus allowing to use a large temporal context without the need to repeatedly processing frames. With this architecture our method obtains state-of-the-art results on the RAI dataset, while running at an unprecedented speed of more than 120x real-time.
Michael Gygli
CBMI1
2018 PHD-GIFs: Personalized Highlight Detection for Automatic GIF Creation
abstract
Highlight detection models are typically trained to identify cues that make visual content appealing or interesting for the general public, with the objective of reducing a video to such moments. However, this "interestingness" of a video segment or image is subjective. Thus, such highlight models provide results of limited relevance for the individual user. On the other hand, training one model per user is inefficient and requires large amounts of personal information which is typically not available. To overcome these limitations, we present a global ranking model which can condition on a particular user's interests. Rather than training one model per user, our model is personalized via its inputs, which allows it to effectively adapt its predictions, given only a few user-specific examples. To train this model, we create a large-scale dataset of users and the GIFs they created, giving us an accurate indication of their interests. Our experiments show that using the user history substantially improves the prediction accuracy. On a test set of 850 videos, our model improves the recall by 8% with respect to generic highlight detectors. Furthermore, our method proves more precise than the user-agnostic baselines even with only one single person-specific example.
Ana Garcia del Molino, Michael Gygli
ACM Multimedia2
2018 AENet: Learning Deep Audio Features for Video Analysis
abstract
We propose a new deep network for audio event recognition, called AENet. In contrast to speech, sounds coming from audio events may be produced by a wide variety of sources. Furthermore, distinguishing them often requires analyzing an extended time period due to the lack of clear subword units that are present in speech. In order to incorporate this long-time frequency structure of audio events, we introduce a convolutional neural network (CNN) operating on a large temporal input. In contrast to previous works, this allows us to train an audio event detection system end to end. The combination of our network architecture and a novel data augmentation outperforms previous methods for audio event detection by 16%. Furthermore, we perform transfer learning and show that our model learned generic audio features, similar to the way CNNs learn generic features on vision tasks. In video analysis, combining visual features and traditional audio features, such as mel frequency cepstral coefficients, typically only leads to marginal improvements. Instead, combining visual features with our AENet features, which can be computed efficiently on a GPU, leads to significant performance improvements on action recognition and video highlight detection. In video highlight detection, our audio features improve the performance by more than 8% over visual features alone.
Naoya Takahashi, Michael Gygli, Luc Van Gool
IEEE Trans. Multim.2
2017 PathTrack: Fast Trajectory Annotation with Path Supervision
abstract
Progress in Multiple Object Tracking (MOT) has been historically limited by the size of the available datasets. We present an efficient framework to annotate trajectories and use it to produce a MOT dataset of unprecedented size. In our novel path supervision the annotator loosely follows the object with the cursor while watching the video, providing a path annotation for each object in the sequence. Our approach is able to turn such weak annotations into dense box trajectories. Our experiments on existing datasets prove that our framework produces more accurate annotations than the state of the art, in a fraction of the time. We further validate our approach by crowdsourcing the PathTrack dataset, with more than 15,000 person trajectories in 720 sequences. Tracking approaches can benefit training on such large-scale datasets, as did object recognition. We prove this by re-training an off-the-shelf person matching network, originally trained on the MOT15 dataset, almost halving the misclassification rate. Additionally, training on our data consistently improves tracking results, both on our dataset and on MOT15. On the latter, we improve the top-performing tracker (NOMT) dropping the number of ID Switches by 18% and fragments by 5%.
Santiago Manen, Michael Gygli, Dengxin Dai, Luc Van Gool
ICCV2
2017 Deep Value Networks Learn to Evaluate and Iteratively Refine Structured Outputs
abstract
We approach structured output prediction by optimizing a deep value network (DVN) to precisely estimate the task loss on different output configurations for a given input. Once the model is trained, we perform inference by gradient descent on the continuous relaxations of the output variables to find outputs with promising scores from the value network. When applied to image segmentation, the value network takes an image and a segmentation mask as inputs and predicts a scalar estimating the intersection over union between the input and ground truth masks. For multi-label classification, the DVN’s objective is to correctly predict the F1 score for any potential label configuration. The DVN framework achieves the state-of-the-art results on multi-label prediction and image segmentation benchmarks.
Michael Gygli, Mohammad Norouzi 0002, Anelia Angelova
ICML1
2017 Query-adaptive Video Summarization via Quality-aware Relevance Estimation
abstract
Although the problem of automatic video summarization has recently received a lot of attention, the problem of creating a video summary that also highlights elements relevant to a search query has been less studied. We address this problem by posing query-relevant summarization as a video frame subset selection problem, which lets us optimise for summaries which are simultaneously diverse, representative of the entire video, and relevant to a text query. We quantify relevance by measuring the distance between frames and queries in a common textual-visual semantic embedding space induced by a neural network. In addition, we extend the model to capture query-independent properties, such as frame quality. We compare our method against previous state of the art on textual-visual embeddings for thumbnail selection and show that our model outperforms them on relevance prediction. Furthermore, we introduce a new dataset, annotated with diversity and query-specific relevance labels. On this dataset, we train and test our complete model for video summarization and show that it outperforms standard baselines such as Maximal Marginal Relevance.
Arun Balajee Vasudevan, Michael Gygli, Anna Volokitin, Luc Van Gool
ACM Multimedia2
2016 Video2GIF: Automatic Generation of Animated GIFs from Video
abstract
We introduce the novel problem of automatically generating animated GIFs from video. GIFs are short looping video with no sound, and a perfect combination between image and video that really capture our attention. GIFs tell a story, express emotion, turn events into humorous moments, and are the new wave of photojournalism. We pose the question: Can we automate the entirely manual and elaborate process of GIF creation by leveraging the plethora of user generated GIF content? We propose a Robust Deep RankNet that, given a video, generates a ranked list of its segments according to their suitability as GIF. We train our model to learn what visual content is often selected for GIFs by using over 100K user generated GIFs and their corresponding video sources. We effectively deal with the noisy web data by proposing a novel adaptive Huber loss in the ranking formulation. We show that our approach is robust to outliers and picks up several patterns that are frequently present in popular animated GIFs. On our new large-scale benchmark dataset, we show the advantage of our approach over several state-of-the-art methods.
Michael Gygli, Yale Song, Liangliang Cao
CVPR1
2016 Predicting When Saliency Maps are Accurate and Eye Fixations Consistent
abstract
Many computational models of visual attention use image features and machine learning techniques to predict eye fixation locations as saliency maps. Recently, the success of Deep Convolutional Neural Networks (DCNNs) for object recognition has opened a new avenue for computational models of visual attention due to the tight link between visual attention and object recognition. In this paper, we show that using features from DCNNs for object recognition we can make predictions that enrich the information provided by saliency models. Namely, we can estimate the reliability of a saliency model from the raw image, which serves as a meta-saliency measure that may be used to select the best saliency algorithm for an image. Analogously, the consistency of the eye fixations among subjects, i.e. the agreement between the eye fixation locations of different subjects, can also be predicted and used by a designer to assess whether subjects reach a consensus about salient image locations.
Anna Volokitin, Michael Gygli, Xavier Boix
CVPR2
2016 Deep Convolutional Neural Networks and Data Augmentation for Acoustic Event Recognition
Naoya Takahashi, Michael Gygli, Beat Pfister, Luc Van Gool
INTERSPEECH2
2016 Analyzing and Predicting GIF Interestingness
abstract
Animated GIFs have regained huge popularity. They are used in instant messaging, online journalism, social media, among others. In this paper, we present an in-depth study on the interestingness of GIFs. We create and annotate a dataset with a set of affective labels, which allows us to investigate the sources of interest. We show that GIFs of pets are considered more interesting that GIFs of people. Furthermore, we study the connection of interest to other features and factors such as popularity. Finally, we build a predictive model and show that it can estimate GIF interestingness with high accuracy. Our model outperforms the existing methods on GIF popularity, as well as a model based on still image interestingness, by a large margin. We envision that the insights and method developed can be used for automatic recognition and generation of interesting GIFs.
Michael Gygli, Mohammad Soleymani 0001
ACM Multimedia1
2015 Video summarization by learning submodular mixtures of objectives
abstract
We present a novel method for summarizing raw, casually captured videos. The objective is to create a short summary that still conveys the story. It should thus be both, interesting and representative for the input video. Previous methods often used simplified assumptions and only optimized for one of these goals. Alternatively, they used handdefined objectives that were optimized sequentially by making consecutive hard decisions. This limits their use to a particular setting. Instead, we introduce a new method that (i) uses a supervised approach in order to learn the importance of global characteristics of a summary and (ii) jointly optimizes for multiple objectives and thus creates summaries that posses multiple properties of a good summary. Experiments on two challenging and very diverse datasets demonstrate the effectiveness of our method, where we outperform or match current state-of-the-art.
Michael Gygli, Helmut Grabner, Luc Van Gool
CVPR1
2014 Creating Summaries from User Videos
Michael Gygli, Helmut Grabner, Hayko Riemenschneider, Luc Van Gool
ECCV (7)1
2013 Sparse Quantization for Patch Description
abstract
The representation of local image patches is crucial for the good performance and efficiency of many vision tasks. Patch descriptors have been designed to generalize towards diverse variations, depending on the application, as well as the desired compromise between accuracy and efficiency. We present a novel formulation of patch description, that serves such issues well. Sparse quantization lies at its heart. This allows for efficient encodings, leading to powerful, novel binary descriptors, yet also to the generalization of existing descriptors like SIFT or BRIEF. We demonstrate the capabilities of our formulation for both key point matching and image classification. Our binary descriptors achieve state-of-the-art results for two key point matching benchmarks, namely those by Brown and Mikolajczyk. For image classification, we propose new descriptors, that perform similar to SIFT on Caltech101 and PASCAL VOC07.
Xavier Boix, Michael Gygli, Gemma Roig, Luc Van Gool
CVPR2
2013 The Interestingness of Images
abstract
We investigate human interest in photos. Based on our own and others' psychophysical experiments, we identify various cues for "interestingness", namely aesthetics, unusualness and general preferences. For the ranking of retrieved images, interestingness shows to be more appropriate than cues proposed earlier. Interestingness is correlated with what people believe they will remember. This is opposed to actual memorability, which is uncorrelated to both. We introduce a set of features computationally capturing the three main aspects of visual interestingness and build an interestingness predictor from them. Its performance is shown on three datasets with varying context, reflecting the prior knowledge of the viewers.
Michael Gygli, Helmut Grabner, Hayko Riemenschneider, Fabian Nater, Luc Van Gool
ICCV1