Marcin Eichner

dblp:25/8610 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Vision and language · 33% Representation and self-supervised learning · 17% 3D vision · 17%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 16 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › vision-language model
multimodal large language model
1.722025
Contrastive Localized Language-Image Pre-Training · ICML 2025
MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs · ICCV 2025
Computer vision › Vision and language
vision-language pretraining
1.722025
Contrastive Localized Language-Image Pre-Training · ICML 2025
Multimodal Autoregressive Pre-training of Large Vision Encoders · CVPR 2025
Computer vision › 3D vision
3d scene understanding
0.912025
MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs · ICCV 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning › spatial reasoning
3d spatial reasoning
0.912025
MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs · ICCV 2025
Machine learning › Representation and self-supervised learning
contrastive learning
0.912025
Contrastive Localized Language-Image Pre-Training · ICML 2025
Machine learning › Representation and self-supervised learning › pre-training
multimodal pretraining
0.912025
Multimodal Autoregressive Pre-training of Large Vision Encoders · CVPR 2025
Computer vision › Image recognition and object detection › object recognition
region-based recognition
0.912025
Contrastive Localized Language-Image Pre-Training · ICML 2025
Computer vision › 3D vision
spatial understanding
0.912025
MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs · ICCV 2025
Computer vision › Face, body and person analysis
human pose estimation
0.542012
Human Pose Co-Estimation and Applications · IEEE Trans. Pattern Anal. Mach. Intell. 2012
2D Articulated Human Pose Estimation and Retrieval in (Almost) Unconstrained Still Images · Int. J. Comput. Vis. 2012
Has My Algorithm Succeeded? An Evaluator for Human Pose Estimators · ECCV (3) 2012
Computer vision › Image recognition and object detection
image classification
0.312025
Multimodal Autoregressive Pre-training of Large Vision Encoders · CVPR 2025
Information retrieval
cross-modal retrieval
0.312025
Contrastive Localized Language-Image Pre-Training · ICML 2025
Computer vision › Face, body and person analysis › gaze analysis
eye contact detection
0.212014
Detecting People Looking at Each Other in Videos · Int. J. Comput. Vis. 2014
Computer vision › Face, body and person analysis
gaze analysis
0.212014
Detecting People Looking at Each Other in Videos · Int. J. Comput. Vis. 2014
Computer vision › Face, body and person analysis
person search
0.112012
2D Articulated Human Pose Estimation and Retrieval in (Almost) Unconstrained Still Images · Int. J. Comput. Vis. 2012
Computer vision › Face, body and person analysis › human pose estimation
multi-person pose estimation
0.112010
We Are Family: Joint Pose Estimation of Multiple Persons · ECCV (1) 2010
Computer vision › Face, body and person analysis › human pose estimation
articulated pose estimation
0.012010
We Are Family: Joint Pose Estimation of Multiple Persons · ECCV (1) 2010

Methods — techniques the papers use, named apart from their topics

region-text contrastive loss · 1.7promptable embeddings · 1.7captioning · 1.7multimodal large language model · 0.9contrastive learning · 0.9autoregressive pretraining · 0.9pictorial structures · 0.1image search · 0.1joint pose estimation · 0.1deep learning · 0.1
YearPublicationVenuePosition
2025 Multimodal Autoregressive Pre-training of Large Vision Encoders
abstract
We introduce a novel method for pre-training of large-scale vision encoders. Building on recent advancements in autoregressive pre-training of vision models, we extend this framework to a multimodal setting, i.e., images and text. In this paper, we present AIMV2, a family of generalist vision encoders characterized by a straightforward pre-training process, scalability, and remarkable performance across a range of downstream tasks. This is achieved by pairing the vision encoder with a multimodal decoder that autoregressively generates raw image patches and text tokens. Our encoders excel not only in multimodal evaluations but also in vision benchmarks such as localization, grounding, and classification. Notably, our AIMV2-3B encoder achieves 89.5% accuracy on ImageNet-1k with a frozen trunk. Furthermore, AIMV2 consistently outperforms state-of-the-art contrastive models (e.g., CLIP, SigLIP) in multimodal image understanding across diverse settings.
Enrico Fini, Mustafa Shukor, Xiujun Li, Philipp Dufter, Michal Klein, David Haldimann, Sai Aitharaju, Victor G. T. da Costa, Louis Béthune, Zhe Gan, Alexander Toshev, Marcin Eichner, Moin Nabi, Yinfei Yang, Joshua M. Susskind, Alaaeldin El-Nouby
CVPR12
2025 MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs
Erik A. Daxberger, Nina Wenzel, David Griffiths, Haiming Gang, Justin Lazarow, Gefen Kohavi, Marcin Eichner, Yinfei Yang, Afshin Dehghan, Peter Grasch
ICCV8
2025 Contrastive Localized Language-Image Pre-Training
abstract
CLIP has been a celebrated method for training vision encoders to generate image/text representations facilitating various applications. Recently, it has been widely adopted as the vision backbone of multimodal large language models (MLLMs). The success of CLIP relies on aligning web-crawled noisy text annotations at image levels. However, such criteria may be insufficient for downstream tasks in need of fine-grained vision representations, especially when understanding region-level is demanding for MLLMs. We improve the localization capability of CLIP with several advances. Our proposed pre-training method, Contrastive Localized Language-Image Pre-training (CLOC), complements CLIP with region-text contrastive loss and modules. We formulate a new concept, promptable embeddings, of which the encoder produces image embeddings easy to transform into region representations given spatial hints. To support large-scale pre-training, we design a visually-enriched and spatially-localized captioning framework to effectively generate region-text labels. By scaling up to billions of annotated images, CLOC enables high-quality regional embeddings for recognition and retrieval tasks, and can be a drop-in replacement of CLIP to enhance MLLMs, especially on referring and grounding tasks.
Hong-You Chen, Zhengfeng Lai, Haotian Zhang 0005, Xinze Wang, Marcin Eichner, Keen You, Bowen Zhang 0002, Yinfei Yang, Zhe Gan
ICML5
2023 Spatial Consistency Loss for Training Multi-Label Classifiers from Single-Label Annotations
abstract
Multi-label image classification is more applicable "in the wild" than single-label classification, as natural images usually contain multiple objects. However, exhaustively annotating images with every object of interest is costly and time-consuming. We train multi-label classifiers from datasets where each image is annotated with a single positive label only. As the presence of all other classes is unknown, we propose an Expected Negative loss that builds a set of expected negative labels in addition to the annotated positives. This set is determined based on prediction consistency, by averaging predictions over consecutive training epochs to build robust targets. Moreover, the ‘crop’ data augmentation leads to additional label noise by cropping out the single annotated object. Our novel spatial consistency loss improves supervision and ensures consistency of the spatial feature maps by maintaining per-class running-average heatmaps for each training image. We use MS-COCO, Pascal VOC, NUS-WIDE and CUB-Birds datasets to demonstrate the gains of the Expected Negative loss in combination with consistency and spatial consistency losses. We also demonstrate improved multi-label classification mAP on ImageNet-1K using the ReaL multi-label validation set.
Thomas Verelst, Paul K. Rubenstein, Marcin Eichner, Tinne Tuytelaars, Maxim Berman
WACV3
2014 Detecting People Looking at Each Other in Videos
Manuel J. Marín-Jiménez, Andrew Zisserman, Marcin Eichner, Vittorio Ferrari
Int. J. Comput. Vis.3
2012 Appearance Sharing for Collective Human Pose Estimation
Marcin Eichner, Vittorio Ferrari
ACCV (1)1
2012 Has My Algorithm Succeeded? An Evaluator for Human Pose Estimators
Nataraj Jammalamadaka, Andrew Zisserman, Marcin Eichner, Vittorio Ferrari, C. V. Jawahar
ECCV (3)3
2012 Video retrieval by mimicking poses
abstract
We describe a method for real time video retrieval where the task is to match the 2D human pose of a query. A user can form a query by (i) interactively controlling a stickman on a web based GUI, (ii) uploading an image of the desired pose, or (iii) using the Kinect and acting out the query himself. The method is scalable and is applied to a dataset of 18 films totaling more than three million frames. The real time performance is achieved by searching for approximate nearest neighbors to the query using a random forest of K-D trees. Apart from the query modalities, we introduce two other areas of novelty. First, we show that pose retrieval can proceed using a low dimensional representation. Second, we show that the precision of the results can be improved substantially by combining the outputs of independent human pose estimation algorithms. The performance of the system is assessed quantitatively over a range of pose queries.
Nataraj Jammalamadaka, Andrew Zisserman, Marcin Eichner, Vittorio Ferrari, C. V. Jawahar
ICMR3
2012 2D Articulated Human Pose Estimation and Retrieval in (Almost) Unconstrained Still Images
Marcin Eichner, Manuel J. Marín-Jiménez, Andrew Zisserman, Vittorio Ferrari
Int. J. Comput. Vis.1
2012 Human Pose Co-Estimation and Applications
abstract
Most existing techniques for articulated Human Pose Estimation (HPE)consider each person independently. Here we tackle the problem in a new setting,coined Human Pose Coestimation (PCE), where multiple people are in a common,but unknown pose. The task of PCE is to estimate their poses jointly and toproduce prototypes characterizing the shared pose. Since the poses of the individual people should be similar to the prototype, PCE has less freedom compared to estimating each pose independently, which simplifies the problem.We demonstrate our PCE technique on two applications. The first is estimating the pose of people performing the same activity synchronously, such as during aerobics, cheerleading, and dancing in a group. We show that PCE improves pose estimation accuracy over estimating each person independently. The second application is learning prototype poses characterizing a pose class directly from an image search engine queried by the class name (e.g., “lotus pose”). We show that PCE leads to better pose estimation in such images, and it learns meaningful prototypes which can be used as priors for pose estimation in novel images.
Marcin Eichner, Vittorio Ferrari
IEEE Trans. Pattern Anal. Mach. Intell.1
2010 We Are Family: Joint Pose Estimation of Multiple Persons
Marcin Eichner, Vittorio Ferrari
ECCV (1)1
2009 Better Appearance Models for Pictorial Structures
abstract
We present a novel approach for estimating body part appearance models for pictorial structures. We learn latent relationships between the appearance of different body parts from annotated images, which then help in estimating better appearance models on novel images. The learned appearance models are general, in that they can be plugged into any pictorial structure engine. In a comprehensive evaluation we demonstrate the bene?ts brought by the new appearance models to an existing articulated human pose estimation algorithm, on hundreds of highly challenging images from the TV series Buffy the vampire slayer and the PASCAL VOC 2008 challenge.
Marcin Eichner, Vittorio Ferrari
BMVC1