Markus Marks

dblp:261/2619 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0001-8016-1637ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Representation and self-supervised learning · 27% 3D vision · 24% Segmentation and scene understanding · 15%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning
1.522025
A Closer Look at Benchmarking Self-supervised Pre-training with Image Classification · Int. J. Comput. Vis. 2025
MABe22: A Multi-Species Multi-Task Benchmark for Learned Representations of Behavior · ICML 2023
Machine learning › Transfer learning and domain adaptation › transferability estimation
transfer evaluation
0.912025
A Closer Look at Benchmarking Self-supervised Pre-training with Image Classification · Int. J. Comput. Vis. 2025
Computer vision › 3D vision
depth estimation
0.812024
Text-Image Alignment for Diffusion-Based Perception · CVPR 2024
Computer vision › Segmentation and scene understanding › semantic segmentation
diffusion-based segmentation
0.812024
Text-Image Alignment for Diffusion-Based Perception · CVPR 2024
Computer vision › Vision and language › cross-modal alignment
image-text alignment
0.812024
Text-Image Alignment for Diffusion-Based Perception · CVPR 2024
Computer vision › 3D vision › depth estimation
monocular depth estimation
0.812024
Text-Image Alignment for Diffusion-Based Perception · CVPR 2024
Computer vision › Image recognition and object detection › object detection
open-vocabulary object detection
0.812024
Text-Image Alignment for Diffusion-Based Perception · CVPR 2024
Computer vision › Segmentation and scene understanding
semantic segmentation
0.812024
Text-Image Alignment for Diffusion-Based Perception · CVPR 2024
Machine learning › Representation and self-supervised learning › representation learning
behavior representation learning
0.712023
MABe22: A Multi-Species Multi-Task Benchmark for Learned Representations of Behavior · ICML 2023
Computer vision › Video understanding and tracking
video representation learning
0.712023
MABe22: A Multi-Species Multi-Task Benchmark for Learned Representations of Behavior · ICML 2023
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning
0.412020
Robust Disentanglement of a Few Factors at a Time · NeurIPS 2020
Computer vision › Image recognition and object detection
object localization
0.312025
Probing the Mid-level Vision Capabilities of Self-Supervised Learning · CVPR 2025
Bioinformatics and computational biology › behavioral analysis
animal behavior analysis
0.212023
MABe22: A Multi-Species Multi-Task Benchmark for Learned Representations of Behavior · ICML 2023

Methods — techniques the papers use, named apart from their topics

self-supervised learning · 1.3pose tracking · 1.3linear probing · 0.9k-nearest neighbors · 0.9fine-tuning · 0.9controlled evaluation · 0.9benchmark protocols · 0.9model personalization · 0.8diffusion model · 0.8caption generation · 0.8
YearPublicationVenuePosition
2026 Diffusion-Based Action Recognition Generalizes to Untrained Domains
abstract
Humans can recognize the same actions despite large context and viewpoint variations, such as differences between species (walking in spiders vs. horses), viewpoints (egocentric vs. third-person), and contexts (real life vs movies). Current deep learning models struggle with such generalization. We propose using features generated by a Vision Diffusion Model (VDM), aggregated via a transformer, to achieve human-like action recognition across these challenging conditions. We find that generalization is enhanced by the use of a model conditioned on earlier timesteps of the diffusion process to highlight semantic information over pixel level details in the extracted features. We experimentally explore the generalization properties of our approach in classifying actions across animal species, across different viewing angles, and different recording contexts. Our model sets a new state-of-the-art across all three generalization benchmarks, bringing machine action recognition closer to human-like robustness. Project page: vision.caltech.edu/actiondiff Code: github.com/frankyaoxiao/ActionDiff
Rogério Guimarães, Frank Xiao, Pietro Perona, Markus Marks
WACV4
2026 SAVeD: Learning to Denoise Low-SNR Video for Improved Downstream Performance
abstract
Low signal-to-noise ratio (SNR) videos—such as those from underwater sonar, ultrasound, and microscopy—pose significant challenges for computer vision models, particularly when paired clean imagery for denoising is unavailable. We present Spatiotemporal Augmentations and denoising in Video for Downstream Tasks (SAVeD), a novel self-supervised method that denoises low-SNR sensor videos using only raw noisy data. By leveraging distinctions between foreground and background motion and exaggerating objects with stronger motion signal, SAVeD enhances foreground object visibility and reduces background and camera noise without requiring clean video. SAVeD has a set of architectural optimizations that lead to faster throughput, training, and inference than existing deep learning methods. We also introduce a new denoising metric, FBD, which indicates foreground-background divergence for detection datasets without requiring clean imagery. Our approach achieves state-of-the-art results for classification, detection, tracking, and counting tasks and it does so with fewer training resource requirements than existing deep-learning-based denoising methods. Project page here, Code: https://github.com/suzanne-stathatos/SAVeD.
Suzanne Stathatos, Michael Hobley, Pietro Perona, Markus Marks
WACV4
2025 Probing the Mid-level Vision Capabilities of Self-Supervised Learning
abstract
Mid-level vision capabilities — such as generic object localization and 3D geometric understanding — are not only fundamental to human vision but are also crucial for many real-world applications of computer vision. These abilities emerge with minimal supervision during the early stages of human visual development. Despite their significance, current self-supervised learning (SSL) approaches are primarily designed and evaluated for high-level recognition tasks, leaving their mid-level vision capabilities largely unexamined.In this study, we introduce a suite of benchmark protocols to systematically assess mid-level vision capabilities and present a comprehensive, controlled evaluation of 22 prominent SSL models across 8 mid-level vision tasks. Our experiments reveal a weak correlation between mid-level and high-level task performance. We also identify several SSL methods with highly imbalanced performance across mid-level and high-level capabilities, as well as some that excel in both. Additionally, we investigate key factors contributing to mid-level vision performance, such as pretraining objectives and network architectures. Our study provides a holistic and timely view of what SSL models have learned, complementing existing research that primarily focuses on high-level vision tasks. We hope our findings guide future SSL research to benchmark models not only on high-level vision tasks but on mid-level as well.
Xuweiyi Chen, Markus Marks, Zezhou Cheng
CVPR2
2025 Learning Keypoints for Multi-Agent Behavior Analysis using Self-Supervision
abstract
The study of social interactions and collective behaviors through multi-agent video analysis is crucial in biology. While self-supervised keypoint discovery has emerged as a promising solution to reduce the need for manual keypoint annotations, existing methods often struggle with videos containing multiple interacting agents, especially those of the same species and color. To address this, we introduce B-KinD-multi, a novel approach that leverages pre-trained video segmentation models to guide keypoint discovery in multi-agent scenarios. This eliminates the need for time-consuming manual annotations on new experimental settings and organisms. Extensive evaluations demonstrate improved keypoint regression and downstream behavioral classification in videos of flies, mice, and rats. Furthermore, our method generalizes well to other species, including ants, bees, and humans, highlighting its potential for broad applications in automated keypoint annotation for multi-agent behavior analysis. Code available under: B-KinD-Multi
Daniel Khalil, Christina Liu, Pietro Perona, Jennifer J. Sun, Markus Marks
WACV5
2025 A Closer Look at Benchmarking Self-supervised Pre-training with Image Classification
abstract
Self-supervised learning (SSL) is a machine learning approach where the data itself provides supervision, eliminating the need for external labels. The model is forced to learn about the data's inherent structure or context by solving a pretext task. With SSL, models can learn from abundant and cheap unlabeled data, significantly reducing the cost of training models where labels are expensive or inaccessible. In Computer Vision, SSL is widely used as pre-training followed by a downstream task, such as supervised transfer, few-shot learning on smaller labeled data sets, and/or unsupervised clustering. Unfortunately, it is infeasible to evaluate SSL methods on all possible downstream tasks and objectively measure the quality of the learned representation. Instead, SSL methods are evaluated using in-domain evaluation protocols, such as fine-tuning, linear probing, and k-nearest neighbors (kNN). However, it is not well understood how well these evaluation protocols estimate the representation quality of a pre-trained model for different downstream tasks under different conditions, such as dataset, metric, and model architecture. In this work, we study how classification-based evaluation protocols for SSL correlate and how well they predict downstream performance on different dataset types. Our study includes eleven common image datasets and 26 models that were pre-trained with different SSL methods or have different model backbones. We find that in-domain linear/kNN probing protocols are, on average, the best general predictors for out-of-domain performance. We further investigate the importance of batch normalization for the various protocols and evaluate how robust correlations are for different kinds of dataset domain shifts. In addition, we challenge assumptions about the relationship between discriminative and generative self-supervised methods, finding that most of their performance differences can be explained by changes to model backbones. Supplementary Information: The online version contains supplementary material available at 10.1007/s11263-025-02402-w.
Markus Marks, Manuel Knott 0001, Neehar Kondapaneni, Elijah Cole, Thijs Defraeye, Fernando Pérez-Cruz, Pietro Perona
Int. J. Comput. Vis.1
2024 Text-Image Alignment for Diffusion-Based Perception
abstract
Diffusion models are generative models with impressive text-to-image synthesis capabilities and have spurred a new wave of creative methods for classical machine learning tasks. However, the best way to harness the perceptual knowledge of these generative models for visual tasks is still an open question. Specifically, it is unclear how to use the prompting interface when applying diffusion backbones to vision tasks. We find that automatically generated captions can improve text-image alignment and significantly enhance a model's cross-attention maps, leading to better perceptual performance. Our approach improves upon the current state-of-the-art (SOTA) in diffusionbased semantic segmentation on ADE20K and the current overall SOTA for depth estimation on NYUv2. Furthermore, our method generalizes to the cross-domain setting. We use model personalization and caption modifications to align our model to the target domain and find improvements over unaligned baselines. Our crossdomain object detection model, trained on Pascal VOC, achieves SOTA results on Watercolor2K. Our cross-domain segmentation method, trained on Cityscapes, achieves SOTA results on Dark Zurich-val and Nighttime Driving. Project page: vision.caltech.edu/TADP/ Code page: github.com/damaggu/TADP
Neehar Kondapaneni, Markus Marks, Manuel Knott 0001, Rogério Guimarães, Pietro Perona
CVPR2
2023 MABe22: A Multi-Species Multi-Task Benchmark for Learned Representations of Behavior
abstract
We introduce MABe22, a large-scale, multi-agent video and trajectory benchmark to assess the quality of learned behavior representations. This dataset is collected from a variety of biology experiments, and includes triplets of interacting mice (4.7 million frames video+pose tracking data, 10 million frames pose only), symbiotic beetle-ant interactions (10 million frames video data), and groups of interacting flies (4.4 million frames of pose tracking data). Accompanying these data, we introduce a panel of real-life downstream analysis tasks to assess the quality of learned representations by evaluating how well they preserve information about the experimental conditions (e.g. strain, time of day, optogenetic stimulation) and animal behavior. We test multiple state-of-the-art self-supervised video and trajectory representation learning methods to demonstrate the use of our benchmark, revealing that methods developed using human action datasets do not fully translate to animal datasets. We hope that our benchmark and dataset encourage a broader exploration of behavior representation learning methods across species and settings.
Jennifer J. Sun, Markus Marks, Andrew Ulmer, Dipam Chakraborty, Brian Geuther, Edward Hayes, Heng Jia, Sebastian Oleszko, Zachary Partridge, Milan Peelman, Alice Robie, Catherine E. Schretter, Keith Sheppard, Param Uttarwar, Julian Morgan Wagner, Erik Werner, Joseph Parker, Pietro Perona, Yisong Yue, Kristin Branson, Ann Kennedy
ICML2
2020 Robust Disentanglement of a Few Factors at a Time
Benjamin Estermann, Markus Marks, Mehmet Fatih Yanik
NeurIPS2