EDBT 2026 Demo / reviewers in the wild / expert
Filip Pavetic
dblp:149/2329
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Deep learning architectures and training · 19% Vision and language · 19% 3D vision · 17% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Environmental and earth informatics · 100% |
Topics — the 16 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training › transformer
vision transformer |
1.3 | 2 | 2023 | Scaling Vision Transformers to 22 Billion Parameters · ICML 2023 FlexiViT: One Model for All Patch Sizes · CVPR 2023 |
Computer vision › Vision and language
image captioning |
1.0 | 2 | 2024 | LocCa: Visual Pretraining with Location-aware Captioners · NeurIPS 2024 On Scaling Up a Multilingual Vision and Language Model · CVPR 2024 |
Computer vision › Image recognition and object detection › object localization
bounding box prediction |
0.8 | 1 | 2024 | LocCa: Visual Pretraining with Location-aware Captioners · NeurIPS 2024 |
Computer vision › Vision and language › vision-language model › vision-language foundation model
multilingual vision-language models |
0.8 | 1 | 2024 | On Scaling Up a Multilingual Vision and Language Model · CVPR 2024 |
Machine learning › Representation and self-supervised learning › pre-training
visual pre-training |
0.8 | 1 | 2024 | LocCa: Visual Pretraining with Location-aware Captioners · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › adaptive computation
adaptive inference |
0.7 | 1 | 2023 | FlexiViT: One Model for All Patch Sizes · CVPR 2023 |
Machine learning › Efficient and distributed learning › large-scale learning
large-scale model training |
0.7 | 1 | 2023 | Scaling Vision Transformers to 22 Billion Parameters · ICML 2023 |
Machine learning › Deep learning architectures and training › transformer › vision transformer
vision transformer scaling |
0.7 | 1 | 2023 | Scaling Vision Transformers to 22 Billion Parameters · ICML 2023 |
Machine learning › Transfer learning and domain adaptation
domain shift |
0.6 | 1 | 2022 | The Auto Arborist Dataset: A Large-Scale Benchmark for Multiview Urban Forest Monitoring Under Domain Shift · CVPR 2022 |
Computer vision › 3D vision › 3d scene modeling › scene representation
neural scene representation |
0.6 | 1 | 2022 | Object Scene Representation Transformer · NeurIPS 2022 |
Computer vision › 3D vision
novel view synthesis |
0.6 | 1 | 2022 | Object Scene Representation Transformer · NeurIPS 2022 |
Computer vision › Segmentation and scene understanding › image decomposition
object-centric scene decomposition |
0.6 | 1 | 2022 | Object Scene Representation Transformer · NeurIPS 2022 |
Computer vision › Image recognition and object detection
object detection |
0.2 | 1 | 2024 | On Scaling Up a Multilingual Vision and Language Model · CVPR 2024 |
Computer vision › Vision and language
visual question answering |
0.2 | 1 | 2024 | On Scaling Up a Multilingual Vision and Language Model · CVPR 2024 |
Machine learning › Trustworthy machine learning › fairness › fairness trade-off
fairness-accuracy trade-off |
0.2 | 1 | 2023 | Scaling Vision Transformers to 22 Billion Parameters · ICML 2023 |
Machine learning › Trustworthy machine learning › fairness
fairness and robustness |
0.2 | 1 | 2023 | Scaling Vision Transformers to 22 Billion Parameters · ICML 2023 |
Methods — techniques the papers use, named apart from their topics
street-level imagery · 1.1multiview imagery · 1.1aerial imagery · 1.1scaling · 0.8multimodal pretraining · 0.8encoder-decoder architecture · 0.8captioning · 0.8patch size randomization · 0.7linear probing on frozen features · 0.7large-scale pretraining · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | On Scaling Up a Multilingual Vision and Language ModelabstractWe explore the boundaries of scaling up a multilingual vision and language model, both in terms of size of the components and the breadth of its training task mixture. Our model achieves new levels of performance on a wide-range of varied and complex tasks, including multiple image-based captioning and question-answering tasks, image-based document understanding and few-shot (in-context) learning, as well as object detection, video question answering, and video captioning. Our model advances the state-of-the-art on most vision-and-language benchmarks considered (20+ of them). Finally, we observe emerging capabilities, such as complex counting and multilingual object detection, tasks that are not explicitly in the training mix. Xi Chen 0071, Josip Djolonga, Piotr Padlewski, Basil Mustafa, Soravit Changpinyo, Carlos Riquelme, Sebastian Goodman, Xiao Wang 0038, Yi Tay, Siamak Shakeri, Mostafa Dehghani 0001, Daniel Salz, Mario Lucic, Michael Tschannen, Arsha Nagrani, Hexiang Hu, Mandar Joshi, Bo Pang 0001, Ceslee Montgomery, Paulina Pietrzyk, Marvin Ritter, A. J. Piergiovanni, Matthias Minderer, Filip Pavetic, Austin Waters, Gang Li 0021, Ibrahim Alabdulmohsin, Lucas Beyer, Julien Amelot, Kenton Lee, Andreas Steiner 0001, Yang Li 0058, Daniel Keysers, Anurag Arnab, Yuanzhong Xu, Keran Rong, Alexander Kolesnikov 0003, Mojtaba Seyedhosseini, Anelia Angelova, Xiaohua Zhai, Neil Houlsby, Radu Soricut |
CVPR | 25 |
| 2024 | LocCa: Visual Pretraining with Location-aware CaptionersabstractImage captioning was recently found to be an effective pretraining method similar to contrastive pretraining. This opens up the largely-unexplored potential of using natural language as a flexible and powerful interface for handling diverse pretraining tasks. In this paper, we demonstrate this with a novel visual pretraining paradigm, LocCa, that incorporates location-aware tasks into captioners to teach models to extract rich information from images. Specifically, LocCa employs two tasks, bounding box prediction and location-dependent captioning, conditioned on the image pixel input. Thanks to the multitask capabilities of an encoder-decoder architecture, we show that an image captioner can effortlessly handle multiple tasks during pretraining. LocCa significantly outperforms standard captioners on downstream localization tasks, achieving state-of-the-art results on RefCOCO/+/g, while maintaining comparable performance on holistic tasks. Our work paves the way for further exploration of natural language interfaces in visual pretraining. Michael Tschannen, Yongqin Xian, Filip Pavetic, Ibrahim Alabdulmohsin, Xiao Wang 0038, André Susano Pinto, Andreas Steiner 0001, Lucas Beyer, Xiaohua Zhai |
NeurIPS | 4 |
| 2023 | FlexiViT: One Model for All Patch SizesabstractVision Transformers convert images to sequences by slicing them into patches. The size of these patches controls a speed/accuracy tradeoff, with smaller patches leading to higher accuracy at greater computational cost, but changing the patch size typically requires retraining the model. In this paper, we demonstrate that simply randomizing the patch size at training time leads to a single set of weights that performs well across a wide range of patch sizes, making it possible to tailor the model to different compute budgets at deployment time. We extensively evaluate the resulting model, which we call FlexiViT, on a wide range of tasks, including classification, image-text retrieval, open-world detection, panoptic segmentation, and semantic segmentation, concluding that it usually matches, and sometimes outperforms, standard ViT models trained at a single patch size in an otherwise identical setup. Hence, FlexiViT training is a simple drop-in improvement for ViT that makes it easy to add compute-adaptive capabilities to most models relying on a ViT backbone architecture. Code and pre-trained models are available at github.com/google-research/big_vision. Lucas Beyer, Pavel Izmailov, Alexander Kolesnikov 0003, Mathilde Caron, Simon Kornblith, Xiaohua Zhai, Matthias Minderer, Michael Tschannen, Ibrahim Alabdulmohsin, Filip Pavetic |
CVPR | 10 |
| 2023 | Scaling Vision Transformers to 22 Billion ParametersabstractThe scaling of Transformers has driven breakthrough capabilities for language models. At present, the largest large language models (LLMs) contain upwards of 100B parameters. Vision Transformers (ViT) have introduced the same architecture to image and video modelling, but these have not yet been successfully scaled to nearly the same degree; the largest dense ViT contains 4B parameters (Chen et al., 2022). We present a recipe for highly efficient and stable training of a 22B-parameter ViT (ViT-22B) and perform a wide variety of experiments on the resulting model. When evaluated on downstream tasks (often with a lightweight linear model on frozen features), ViT-22B demonstrates increasing performance with scale. We further observe other interesting benefits of scale, including an improved tradeoff between fairness and performance, state-of-the-art alignment to human visual perception in terms of shape/texture bias, and improved robustness. ViT-22B demonstrates the potential for "LLM-like" scaling in vision, and provides key steps towards getting there. Mostafa Dehghani 0001, Josip Djolonga, Basil Mustafa, Piotr Padlewski, Jonathan Heek, Justin Gilmer, Andreas Steiner 0001, Mathilde Caron, Robert Geirhos, Ibrahim Alabdulmohsin, Rodolphe Jenatton, Lucas Beyer, Michael Tschannen, Anurag Arnab, Xiao Wang 0038, Carlos Riquelme, Matthias Minderer, Joan Puigcerver, Utku Evci, Sjoerd van Steenkiste, Gamaleldin F. Elsayed, Aravindh Mahendran, Fisher Yu 0001, Avital Oliver, Fantine Huot, Jasmijn Bastings, Mark Collier, Alexey A. Gritsenko, Vighnesh Birodkar, Cristina Nader Vasconcelos, Yi Tay, Thomas Mensink, Alexander Kolesnikov 0003, Filip Pavetic, Dustin Tran, Thomas Kipf, Mario Lucic, Xiaohua Zhai, Daniel Keysers, Jeremiah J. Harmsen, Neil Houlsby |
ICML | 35 |
| 2022 | The Auto Arborist Dataset: A Large-Scale Benchmark for Multiview Urban Forest Monitoring Under Domain ShiftabstractGeneralization to novel domains is a fundamental chal-lenge for computer vision. Near-perfect accuracy on bench-marks is common, but these models do not work as expected when deployed outside of the training distribution. To build computer vision systems that truly solve real-world prob-lems at global scale, we need benchmarks that fully capture real-world complexity, including geographic domain shift, long-tailed distributions, and data noise. We propose urban forest monitoring as an ideal testbed for studying and improving upon these computer vision challenges, while working towards filling a crucial environ-mental and societal need. Urban forests provide significant benefits to urban societies. However, planning and main-taining these forests is expensive. One particularly costly aspect of urban forest management is monitoring the ex-isting trees in a city: e.g., tracking tree locations, species, and health. Monitoring efforts are currently based on tree censuses built by human experts, costing cities millions of dollars per census and thus collected infrequently. Previous investigations into automating urban forest monitoring focused on small datasets from single cities, covering only common categories. To address these short-comings, we introduce a new large-scale dataset that joins public tree censuses from 23 cities with a large collection of street level and aerial imagery. Our Auto Arborist dataset contains over 2.5M trees and 344 genera and is >2 or-ders of magnitude larger than the closest dataset in the literature. We introduce baseline results on our dataset across modalities as well as metrics for the detailed analy-sis of generalization with respect to geographic distribution shifts, vital for such a system to be deployed at-scale. Sara Beery, Guanhang Wu, Trevor Edwards, Filip Pavetic, Bo Majewski, Shreyasee Mukherjee, Stanley Chan, John Morgan, Vivek Rathod, Jonathan Huang |
CVPR | 4 |
| 2022 | Object Scene Representation TransformerabstractA compositional understanding of the world in terms of objects and their geometry in 3D space is considered a cornerstone of human cognition. Facilitating the learning of such a representation in neural networks holds promise for substantially improving labeled data efficiency. As a key step in this direction, we make progress on the problem of learning 3D-consistent decompositions of complex scenes into individual objects in an unsupervised fashion. We introduce Object Scene Representation Transformer (OSRT), a 3D-centric model in which individual object representations naturally emerge through novel view synthesis. OSRT scales to significantly more complex scenes with larger diversity of objects and backgrounds than existing methods. At the same time, it is multiple orders of magnitude faster at compositional rendering thanks to its light field parametrization and the novel Slot Mixer decoder. We believe this work will not only accelerate future architecture exploration and scaling efforts, but it will also serve as a useful tool for both object-centric as well as neural scene representation learning communities. Mehdi S. M. Sajjadi, Daniel Duckworth, Aravindh Mahendran, Sjoerd van Steenkiste, Filip Pavetic, Mario Lucic, Leonidas J. Guibas, Klaus Greff, Thomas Kipf |
NeurIPS | 5 |