Tobias Weyand

dblp:71/6931 · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 6 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Representation and self-supervised learning · 38% Video understanding and tracking · 24% Image recognition and object detection · 18%
Computer graphics and multimedia
3 papers
Multimedia analysis and retrieval · 100%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%

Topics — the 20 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
masked autoencoder
1.522024
Extending Video Masked Autoencoders to 128 frames · NeurIPS 2024
VideoPrism: A Foundational Visual Encoder for Video Understanding · ICML 2024
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked autoencoder
video masked autoencoder
1.522024
Extending Video Masked Autoencoders to 128 frames · NeurIPS 2024
VideoPrism: A Foundational Visual Encoder for Video Understanding · ICML 2024
Computer vision › Video understanding and tracking
video representation learning
0.812024
VideoPrism: A Foundational Visual Encoder for Video Understanding · ICML 2024
Computer vision › 3D vision › visual localization › geo-localization
image geo-localization
0.622018
CPlaNet: Enhancing Image Geolocalization by Combinatorial Partitioning of Maps · ECCV (10) 2018
PlaNet - Photo Geolocation with Convolutional Neural Networks · ECCV (8) 2016
Computer vision › Image recognition and object detection › food recognition
food nutrition estimation
0.512021
Nutrition5k: Towards Automatic Nutritional Understanding of Generic Food · CVPR 2021
Computer vision › Image recognition and object detection
food recognition
0.512021
Nutrition5k: Towards Automatic Nutritional Understanding of Generic Food · CVPR 2021
Medical and health informatics
dietary assessment
0.512021
Nutrition5k: Towards Automatic Nutritional Understanding of Generic Food · CVPR 2021
Computer vision › Image recognition and object detection › image classification
fine-grained image classification
0.412020
Google Landmarks Dataset v2 - A Large-Scale Benchmark for Instance-Level Recognition and Retrieval · CVPR 2020
Multimedia analysis and retrieval
image retrieval
0.412020
Google Landmarks Dataset v2 - A Large-Scale Benchmark for Instance-Level Recognition and Retrieval · CVPR 2020
Multimedia analysis and retrieval › image retrieval
instance-level image retrieval
0.412020
Google Landmarks Dataset v2 - A Large-Scale Benchmark for Instance-Level Recognition and Retrieval · CVPR 2020
Multimedia analysis and retrieval › multimedia analysis › multimedia collection analysis
image collection analysis
0.322013
Discovering Details and Scene Structure with Hierarchical Iconoid Shift · ICCV 2013
Discovering favorite views of popular places with iconoid shift · ICCV 2011
Computer vision › 3D vision
local feature descriptor
0.312017
Large-Scale Image Retrieval with Attentive Deep Local Features · ICCV 2017
Information retrieval
image retrieval
0.312017
Large-Scale Image Retrieval with Attentive Deep Local Features · ICCV 2017
Information retrieval › image retrieval
large-scale image retrieval
0.312017
Large-Scale Image Retrieval with Attentive Deep Local Features · ICCV 2017
Computer vision › Video understanding and tracking
video question answering
0.312025
Minerva: Evaluating Complex Video Reasoning · ICCV 2025
Machine learning › Efficient and distributed learning › token reduction
token downsampling
0.212024
Extending Video Masked Autoencoders to 128 frames · NeurIPS 2024
Computer vision › Vision and language › vision-language pretraining
video-language pre-training
0.212024
VideoPrism: A Foundational Visual Encoder for Video Understanding · ICML 2024
Multimedia analysis and retrieval › object recognition
landmark recognition
0.012013
Discovering Details and Scene Structure with Hierarchical Iconoid Shift · ICCV 2013
Computer vision › 3D vision › multi-view geometry › two-view geometry
homography
0.012011
Discovering favorite views of popular places with iconoid shift · ICCV 2011
Computer vision › 3D vision
multi-view geometry
0.012011
Discovering favorite views of popular places with iconoid shift · ICCV 2011

Methods — techniques the papers use, named apart from their topics

depth sensing · 1.0computer vision regression · 1.0convolutional neural network · 0.8tokenizer · 0.8token shuffling · 0.8masked autoencoding · 0.8masked autoencoder · 0.8global-local distillation · 0.8adaptive masking · 0.8attention mechanism · 0.6hierarchical medoid shift · 0.2dendrogram construction · 0.2medoid shift · 0.1matching graph exploration · 0.1
YearPublicationVenuePosition
2025 Minerva: Evaluating Complex Video Reasoning
Arsha Nagrani, Sachit Menon, Ahmet Iscen, Shyamal Buch, Ramin Mehran, Nilpa Jha, Anja Hauth, Yukun Zhu, Carl Vondrick, Mikhail Sirotenko, Cordelia Schmid, Tobias Weyand
ICCV12
2024 VideoPrism: A Foundational Visual Encoder for Video Understanding
abstract
We introduce VideoPrism, a general-purpose video encoder that tackles diverse video understanding tasks with a single frozen model. We pretrain VideoPrism on a heterogeneous corpus containing 36M high-quality video-caption pairs and 582M video clips with noisy parallel text (e.g., ASR transcripts). The pretraining approach improves upon masked autoencoding by global-local distillation of semantic video embeddings and a token shuffling scheme, enabling VideoPrism to focus primarily on the video modality while leveraging the invaluable text associated with videos. We extensively test VideoPrism on four broad groups of video understanding tasks, from web video question answering to CV for science, achieving state-of-the-art performance on 31 out of 33 video understanding benchmarks.
Long Zhao 0003, Nitesh Bharadwaj Gundavarapu, Liangzhe Yuan, Hao Zhou 0014, Shen Yan 0008, Jennifer J. Sun, Luke Friedman, Rui Qian 0003, Tobias Weyand, Yue Zhao 0006, Rachel Hornung, Florian Schroff, Ming-Hsuan Yang 0001, David A. Ross, Huisheng Wang, Hartwig Adam, Mikhail Sirotenko, Ting Liu 0005, Boqing Gong
ICML9
2024 Extending Video Masked Autoencoders to 128 frames
abstract
Video understanding has witnessed significant progress with recent video foundation models demonstrating strong performance owing to self-supervised pre-training objectives; Masked Autoencoders (MAE) being the design of choice. Nevertheless, the majority of prior works that leverage MAE pre-training have focused on relatively short video representations (16 / 32 frames in length) largely due to hardware memory and compute limitations that scale poorly with video length due to the dense memory-intensive self-attention decoding. One natural strategy to address these challenges is to subsample tokens to reconstruct during decoding (or decoder masking). In this work, we propose an effective strategy for prioritizing tokens which allows training on longer video sequences (128 frames) and gets better performance than, more typical, random and uniform masking strategies. The core of our approach is an adaptive decoder masking strategy that prioritizes the most important tokens and uses quantized tokens as reconstruction objectives. Our adaptive strategy leverages a powerful MAGVIT-based tokenizer that jointly learns the tokens and their priority. We validate our design choices through exhaustive ablations and observe improved performance of the resulting long-video (128 frames) encoders over short-video (32 frames) counterparts. With our long-video masked autoencoder (LVMAE) strategy, we surpass state-of-the-art on Diving48 by 3.9 points and EPIC-Kitchens-100 verb classification by 2.5 points while relying on a simple core architecture and video-only pre-training (unlike some of the prior works that require millions of labeled video-text pairs or specialized encoders).
Nitesh Bharadwaj Gundavarapu, Luke Friedman, Raghav Goyal, Chaitra Hegde, Eirikur Agustsson, Sagar Waghmare, Mikhail Sirotenko, Ming-Hsuan Yang 0001, Tobias Weyand, Boqing Gong, Leonid Sigal
NeurIPS9
2021 Nutrition5k: Towards Automatic Nutritional Understanding of Generic Food
abstract
Understanding the nutritional content of food from visual data is a challenging computer vision problem, with the potential to have a positive and widespread impact on public health. Studies in this area are limited to existing datasets in the field that lack sufficient diversity or labels required for training models with nutritional understanding capability. We introduce Nutrition5k, a novel dataset of 5k diverse, real world food dishes with corresponding video streams, depth images, component weights, and high accuracy nutritional content annotation. We demonstrate the potential of this dataset by training a computer vision algorithm capable of predicting the caloric and macronutrient values of a complex, real world dish at an accuracy that outperforms professional nutritionists. Further we present a baseline for incorporating depth sensor data to improve nutrition predictions. We release Nutrition5k in the hope that it will accelerate innovation in the space of nutritional understanding. The dataset is available at https://github.com/google-research-datasets/Nutrition5k.
Quin Thames, Arjun Karpur, Wade Norris, Fangting Xia, Liviu Panait, Tobias Weyand, Jack Sim
CVPR6
2020 Google Landmarks Dataset v2 - A Large-Scale Benchmark for Instance-Level Recognition and Retrieval
abstract
While image retrieval and instance recognition techniques are progressing rapidly, there is a need for challenging datasets to accurately measure their performance -- while posing novel challenges that are relevant for practical applications. We introduce the Google Landmarks Dataset v2 (GLDv2), a new benchmark for large-scale, fine-grained instance recognition and image retrieval in the domain of human-made and natural landmarks. GLDv2 is the largest such dataset to date by a large margin, including over 5M images and 200k distinct instance labels. Its test set consists of 118k images with ground truth annotations for both the retrieval and recognition tasks. The ground truth construction involved over 800 hours of human annotator work. Our new dataset has several challenging properties inspired by real-world applications that previous datasets did not consider: An extremely long-tailed class distribution, a large fraction of out-of-domain test photos and large intra-class variability. The dataset is sourced from Wikimedia Commons, the world's largest crowdsourced collection of landmark photos. We provide baseline results for both recognition and retrieval tasks based on state-of-the-art methods as well as competitive results from a public challenge. We further demonstrate the suitability of the dataset for transfer learning by showing that image embeddings trained on it achieve competitive retrieval performance on independent datasets. The dataset images, ground-truth and metric scoring code are available at https://github.com/cvdfoundation/google-landmark.
Tobias Weyand, André Araújo 0001, Bingyi Cao, Jack Sim
CVPR1
2018 CPlaNet: Enhancing Image Geolocalization by Combinatorial Partitioning of Maps
Hongsuck Seo, Tobias Weyand, Jack Sim, Bohyung Han
ECCV (10)2
2017 Large-Scale Image Retrieval with Attentive Deep Local Features
abstract
We propose an attentive local feature descriptor suitable for large-scale image retrieval, referred to as DELE (DEep Local Feature). The new feature is based on convolutional neural networks, which are trained only with image-level annotations on a landmark image dataset. To identify semantically useful local features for image retrieval, we also propose an attention mechanism for key point selection, which shares most network layers with the descriptor. This frame-work can be used for image retrieval as a drop-in replacement for other keypoint detectors and descriptors, enabling more accurate feature matching and geometric verification. Our system produces reliable confidence scores to reject false positives–in particular, it is robust against queries that have no correct match in the database. To evaluate the proposed descriptor, we introduce a new large-scale dataset, referred to as Google-Landmarks dataset, which involves challenges in both database and query such as background clutter, partial occlusion, multiple landmarks, objects in variable scales, etc. We show that DELE outperforms the state-of-the-art global and local descriptors in the large-scale setting by significant margins.
Hyeonwoo Noh, André Araújo 0001, Jack Sim, Tobias Weyand, Bohyung Han
ICCV4
2016 PlaNet - Photo Geolocation with Convolutional Neural Networks
Tobias Weyand, Ilya Kostrikov, James Philbin
ECCV (8)1
2015 Fixing WTFs: Detecting Image Matches Caused by Watermarks, Timestamps, and Frames in Internet Photos
abstract
An increasing number of photos in Internet photo collections comes with watermarks, timestamps, or frames (in the following called WTFs) embedded in the image content. In image retrieval, such WTFs often cause false-positive matches. In image clustering, these false-positive matches can cause clusters of different buildings to be joined into one. This harms applications like landmark recognition or large-scale structure-from-motion, which rely on clean building clusters. We propose a simple, but highly effective detector for such false-positive matches. Given a matching image pair with an estimated homography, we first determine similar regions in both images. Exploiting the fact that WTFs typically appear near the border, we build a spatial histogram of the similar regions and apply a binary classifier to decide whether the match is due to a WTF. Based on a large-scale dataset of WTFs we collected from Internet photo collections, we show that our approach is general enough to recognize a large variety of watermarks, timestamps, and frames, and that it is efficient enough for large scale applications. In addition, we show that our method fixes the problems that WTFs cause in image clustering applications. The source code is publicly available1and easy to integrate into existing retrieval and clustering systems.
Tobias Weyand, Chih-Yun Tsai, Bastian Leibe
WACV1
2015 Visual landmark recognition from Internet photo collections: A large-scale evaluation
Tobias Weyand, Bastian Leibe
Comput. Vis. Image Underst.1
2013 Discovering Details and Scene Structure with Hierarchical Iconoid Shift
abstract
Current landmark recognition engines are typically aimed at recognizing building-scale landmarks, but miss interesting details like portals, statues or windows. This is because they use a flat clustering that summarizes all photos of a building facade in one cluster. We propose Hierarchical Iconoid Shift, a novel landmark clustering algorithm capable of discovering such details. Instead of just a collection of clusters, the output of HIS is a set of dendrograms describing the detail hierarchy of a landmark. HIS is based on the novel Hierarchical Medoid Shift clustering algorithm that performs a continuous mode search over the complete scale space. HMS is completely parameter-free, has the same complexity as Medoid Shift and is easy to parallelize. We evaluate HIS on 800k images of 34 landmarks and show that it can extract an often surprising amount of detail and structure that can be applied, e.g., to provide a mobile user with more detailed information on a landmark or even to extend the landmark's Wikipedia article.
Tobias Weyand, Bastian Leibe
ICCV1
2012 Image Retrieval for Image-Based Localization Revisited
abstract
To reliably determine the camera pose of an image relative to a 3D point cloud of a scene, correspondences between 2D features and 3D points are needed. Recent work has demonstrated that directly matching the features against the points outperforms methods that take an intermediate image retrieval step in terms of the number of images that can be localized successfully. Yet, direct matching is inherently less scalable than retrieval-based approaches. In this paper, we therefore analyze the algorithmic factors that cause the performance gap and identify false positive votes as the main source of the gap. Based on a detailed experimental evaluation, we show that retrieval methods using a selective voting scheme are able to outperform state-of-the-art direct matching methods. We explore how both selective voting and correspondence computation can be accelerated by using a Hamming embedding of feature descriptors. Furthermore, we introduce a new dataset with challenging query images for the evaluation of image-based localization.
Torsten Sattler, Tobias Weyand, Bastian Leibe, Leif Kobbelt
BMVC2
2011 Discovering favorite views of popular places with iconoid shift
abstract
In this paper, we propose a novel algorithm for automatic landmark building discovery in large, unstructured image collections. In contrast to other approaches which aim at a hard clustering, we regard the task as a mode estimation problem. Our algorithm searches for local attractors in the image distribution that have a maximal mutual homography overlap with the images in their neighborhood. Those attractors correspond to central, iconic views of single objects or buildings, which we efficiently extract using a medoid shift search with a novel distance measure. We propose efficient algorithms for performing this search. Most importantly, our approach performs only an efficient local exploration of the matching graph that makes it applicable for large-scale analysis of photo collections. We show experimental results validating our approach on a dataset of 500k images of the inner city of Paris.
Tobias Weyand, Bastian Leibe
ICCV1
2009 Log-Linear Mixtures for Object Class Recognition
abstract
We present log-linear mixture models as a fully discriminative approach to object cate-gory recognition which can, analogously to kernelised models, represent non-linear de-cision boundaries. We show that this model is the discriminative counterpart to Gaussian mixtures and that either one can be transformed into the respective other. However, the proposed model is easier to extend toward fusing multiple cues and numerically more stable to train and to evaluate. Experiments on the PASCAL VOC 2006 data show that the performance of our model compares favourably well to the state-of-the-art despite the model consisting of an order of magnitude fewer parameters, which suggests excellent generalisation capabilities. 1
Tobias Weyand, Thomas Deselaers, Hermann Ney
BMVC1