VLDB 2026 Research / reviewers in the wild / expert
Marcin Eichner
dblp:25/8610
· DBLP profile ↗
12ranked-venue papers
5as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Vision and language · 33% Representation and self-supervised learning · 17% 3D vision · 17% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 16 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › vision-language model
multimodal large language model |
1.7 | 2 | 2025 | Contrastive Localized Language-Image Pre-Training · ICML 2025 MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs · ICCV 2025 |
Computer vision › Vision and language
vision-language pretraining |
1.7 | 2 | 2025 | Contrastive Localized Language-Image Pre-Training · ICML 2025 Multimodal Autoregressive Pre-training of Large Vision Encoders · CVPR 2025 |
Computer vision › 3D vision
3d scene understanding |
0.9 | 1 | 2025 | MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs · ICCV 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › spatial reasoning
3d spatial reasoning |
0.9 | 1 | 2025 | MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs · ICCV 2025 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.9 | 1 | 2025 | Contrastive Localized Language-Image Pre-Training · ICML 2025 |
Machine learning › Representation and self-supervised learning › pre-training
multimodal pretraining |
0.9 | 1 | 2025 | Multimodal Autoregressive Pre-training of Large Vision Encoders · CVPR 2025 |
Computer vision › Image recognition and object detection › object recognition
region-based recognition |
0.9 | 1 | 2025 | Contrastive Localized Language-Image Pre-Training · ICML 2025 |
Computer vision › 3D vision
spatial understanding |
0.9 | 1 | 2025 | MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs · ICCV 2025 |
Computer vision › Face, body and person analysis
human pose estimation |
0.5 | 4 | 2012 | Human Pose Co-Estimation and Applications · IEEE Trans. Pattern Anal. Mach. Intell. 2012 2D Articulated Human Pose Estimation and Retrieval in (Almost) Unconstrained Still Images · Int. J. Comput. Vis. 2012 Has My Algorithm Succeeded? An Evaluator for Human Pose Estimators · ECCV (3) 2012 |
Computer vision › Image recognition and object detection
image classification |
0.3 | 1 | 2025 | Multimodal Autoregressive Pre-training of Large Vision Encoders · CVPR 2025 |
Information retrieval
cross-modal retrieval |
0.3 | 1 | 2025 | Contrastive Localized Language-Image Pre-Training · ICML 2025 |
Computer vision › Face, body and person analysis › gaze analysis
eye contact detection |
0.2 | 1 | 2014 | Detecting People Looking at Each Other in Videos · Int. J. Comput. Vis. 2014 |
Computer vision › Face, body and person analysis
gaze analysis |
0.2 | 1 | 2014 | Detecting People Looking at Each Other in Videos · Int. J. Comput. Vis. 2014 |
Computer vision › Face, body and person analysis
person search |
0.1 | 1 | 2012 | 2D Articulated Human Pose Estimation and Retrieval in (Almost) Unconstrained Still Images · Int. J. Comput. Vis. 2012 |
Computer vision › Face, body and person analysis › human pose estimation
multi-person pose estimation |
0.1 | 1 | 2010 | We Are Family: Joint Pose Estimation of Multiple Persons · ECCV (1) 2010 |
Computer vision › Face, body and person analysis › human pose estimation
articulated pose estimation |
0.0 | 1 | 2010 | We Are Family: Joint Pose Estimation of Multiple Persons · ECCV (1) 2010 |
Methods — techniques the papers use, named apart from their topics
region-text contrastive loss · 1.7promptable embeddings · 1.7captioning · 1.7multimodal large language model · 0.9contrastive learning · 0.9autoregressive pretraining · 0.9pictorial structures · 0.1image search · 0.1joint pose estimation · 0.1deep learning · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multimodal Autoregressive Pre-training of Large Vision EncodersabstractWe introduce a novel method for pre-training of large-scale vision encoders. Building on recent advancements in autoregressive pre-training of vision models, we extend this framework to a multimodal setting, i.e., images and text. In this paper, we present AIMV2, a family of generalist vision encoders characterized by a straightforward pre-training process, scalability, and remarkable performance across a range of downstream tasks. This is achieved by pairing the vision encoder with a multimodal decoder that autoregressively generates raw image patches and text tokens. Our encoders excel not only in multimodal evaluations but also in vision benchmarks such as localization, grounding, and classification. Notably, our AIMV2-3B encoder achieves 89.5% accuracy on ImageNet-1k with a frozen trunk. Furthermore, AIMV2 consistently outperforms state-of-the-art contrastive models (e.g., CLIP, SigLIP) in multimodal image understanding across diverse settings. Enrico Fini, Mustafa Shukor, Xiujun Li, Philipp Dufter, Michal Klein, David Haldimann, Sai Aitharaju, Victor G. T. da Costa, Louis Béthune, Zhe Gan, Alexander Toshev, Marcin Eichner, Moin Nabi, Yinfei Yang, Joshua M. Susskind, Alaaeldin El-Nouby |
CVPR | 12 |
| 2025 | MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs
Erik A. Daxberger, Nina Wenzel, David Griffiths, Haiming Gang, Justin Lazarow, Gefen Kohavi, Marcin Eichner, Yinfei Yang, Afshin Dehghan, Peter Grasch |
ICCV | 8 |
| 2025 | Contrastive Localized Language-Image Pre-TrainingabstractCLIP has been a celebrated method for training vision encoders to generate image/text representations facilitating various applications. Recently, it has been widely adopted as the vision backbone of multimodal large language models (MLLMs). The success of CLIP relies on aligning web-crawled noisy text annotations at image levels. However, such criteria may be insufficient for downstream tasks in need of fine-grained vision representations, especially when understanding region-level is demanding for MLLMs. We improve the localization capability of CLIP with several advances. Our proposed pre-training method, Contrastive Localized Language-Image Pre-training (CLOC), complements CLIP with region-text contrastive loss and modules. We formulate a new concept, promptable embeddings, of which the encoder produces image embeddings easy to transform into region representations given spatial hints. To support large-scale pre-training, we design a visually-enriched and spatially-localized captioning framework to effectively generate region-text labels. By scaling up to billions of annotated images, CLOC enables high-quality regional embeddings for recognition and retrieval tasks, and can be a drop-in replacement of CLIP to enhance MLLMs, especially on referring and grounding tasks. Hong-You Chen, Zhengfeng Lai, Haotian Zhang 0005, Xinze Wang, Marcin Eichner, Keen You, Bowen Zhang 0002, Yinfei Yang, Zhe Gan |
ICML | 5 |
| 2023 | Spatial Consistency Loss for Training Multi-Label Classifiers from Single-Label AnnotationsabstractMulti-label image classification is more applicable "in the wild" than single-label classification, as natural images usually contain multiple objects. However, exhaustively annotating images with every object of interest is costly and time-consuming. We train multi-label classifiers from datasets where each image is annotated with a single positive label only. As the presence of all other classes is unknown, we propose an Expected Negative loss that builds a set of expected negative labels in addition to the annotated positives. This set is determined based on prediction consistency, by averaging predictions over consecutive training epochs to build robust targets. Moreover, the ‘crop’ data augmentation leads to additional label noise by cropping out the single annotated object. Our novel spatial consistency loss improves supervision and ensures consistency of the spatial feature maps by maintaining per-class running-average heatmaps for each training image. We use MS-COCO, Pascal VOC, NUS-WIDE and CUB-Birds datasets to demonstrate the gains of the Expected Negative loss in combination with consistency and spatial consistency losses. We also demonstrate improved multi-label classification mAP on ImageNet-1K using the ReaL multi-label validation set. Thomas Verelst, Paul K. Rubenstein, Marcin Eichner, Tinne Tuytelaars, Maxim Berman |
WACV | 3 |
| 2014 | Detecting People Looking at Each Other in Videos
Manuel J. Marín-Jiménez, Andrew Zisserman, Marcin Eichner, Vittorio Ferrari |
Int. J. Comput. Vis. | 3 |
| 2012 | Appearance Sharing for Collective Human Pose Estimation
Marcin Eichner, Vittorio Ferrari |
ACCV (1) | 1 |
| 2012 | Has My Algorithm Succeeded? An Evaluator for Human Pose Estimators
Nataraj Jammalamadaka, Andrew Zisserman, Marcin Eichner, Vittorio Ferrari, C. V. Jawahar |
ECCV (3) | 3 |
| 2012 | Video retrieval by mimicking posesabstractWe describe a method for real time video retrieval where the task is to match the 2D human pose of a query. A user can form a query by (i) interactively controlling a stickman on a web based GUI, (ii) uploading an image of the desired pose, or (iii) using the Kinect and acting out the query himself. The method is scalable and is applied to a dataset of 18 films totaling more than three million frames. The real time performance is achieved by searching for approximate nearest neighbors to the query using a random forest of K-D trees. Apart from the query modalities, we introduce two other areas of novelty. First, we show that pose retrieval can proceed using a low dimensional representation. Second, we show that the precision of the results can be improved substantially by combining the outputs of independent human pose estimation algorithms. The performance of the system is assessed quantitatively over a range of pose queries. Nataraj Jammalamadaka, Andrew Zisserman, Marcin Eichner, Vittorio Ferrari, C. V. Jawahar |
ICMR | 3 |
| 2012 | 2D Articulated Human Pose Estimation and Retrieval in (Almost) Unconstrained Still Images
Marcin Eichner, Manuel J. Marín-Jiménez, Andrew Zisserman, Vittorio Ferrari |
Int. J. Comput. Vis. | 1 |
| 2012 | Human Pose Co-Estimation and ApplicationsabstractMost existing techniques for articulated Human Pose Estimation (HPE)consider each person independently. Here we tackle the problem in a new setting,coined Human Pose Coestimation (PCE), where multiple people are in a common,but unknown pose. The task of PCE is to estimate their poses jointly and toproduce prototypes characterizing the shared pose. Since the poses of the individual people should be similar to the prototype, PCE has less freedom compared to estimating each pose independently, which simplifies the problem.We demonstrate our PCE technique on two applications. The first is estimating the pose of people performing the same activity synchronously, such as during aerobics, cheerleading, and dancing in a group. We show that PCE improves pose estimation accuracy over estimating each person independently. The second application is learning prototype poses characterizing a pose class directly from an image search engine queried by the class name (e.g., “lotus pose”). We show that PCE leads to better pose estimation in such images, and it learns meaningful prototypes which can be used as priors for pose estimation in novel images. Marcin Eichner, Vittorio Ferrari |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2010 | We Are Family: Joint Pose Estimation of Multiple Persons
Marcin Eichner, Vittorio Ferrari |
ECCV (1) | 1 |
| 2009 | Better Appearance Models for Pictorial StructuresabstractWe present a novel approach for estimating body part appearance models for pictorial structures. We learn latent relationships between the appearance of different body parts from annotated images, which then help in estimating better appearance models on novel images. The learned appearance models are general, in that they can be plugged into any pictorial structure engine. In a comprehensive evaluation we demonstrate the bene?ts brought by the new appearance models to an existing articulated human pose estimation algorithm, on hundreds of highly challenging images from the TV series Buffy the vampire slayer and the PASCAL VOC 2008 challenge. Marcin Eichner, Vittorio Ferrari |
BMVC | 1 |