VLDB 2026 Research / reviewers in the wild / expert
Markus Marks
dblp:261/2619
· DBLP profile ↗
8ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0001-8016-1637ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Representation and self-supervised learning · 27% 3D vision · 24% Segmentation and scene understanding · 15% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 13 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning |
1.5 | 2 | 2025 | A Closer Look at Benchmarking Self-supervised Pre-training with Image Classification · Int. J. Comput. Vis. 2025 MABe22: A Multi-Species Multi-Task Benchmark for Learned Representations of Behavior · ICML 2023 |
Machine learning › Transfer learning and domain adaptation › transferability estimation
transfer evaluation |
0.9 | 1 | 2025 | A Closer Look at Benchmarking Self-supervised Pre-training with Image Classification · Int. J. Comput. Vis. 2025 |
Computer vision › 3D vision
depth estimation |
0.8 | 1 | 2024 | Text-Image Alignment for Diffusion-Based Perception · CVPR 2024 |
Computer vision › Segmentation and scene understanding › semantic segmentation
diffusion-based segmentation |
0.8 | 1 | 2024 | Text-Image Alignment for Diffusion-Based Perception · CVPR 2024 |
Computer vision › Vision and language › cross-modal alignment
image-text alignment |
0.8 | 1 | 2024 | Text-Image Alignment for Diffusion-Based Perception · CVPR 2024 |
Computer vision › 3D vision › depth estimation
monocular depth estimation |
0.8 | 1 | 2024 | Text-Image Alignment for Diffusion-Based Perception · CVPR 2024 |
Computer vision › Image recognition and object detection › object detection
open-vocabulary object detection |
0.8 | 1 | 2024 | Text-Image Alignment for Diffusion-Based Perception · CVPR 2024 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.8 | 1 | 2024 | Text-Image Alignment for Diffusion-Based Perception · CVPR 2024 |
Machine learning › Representation and self-supervised learning › representation learning
behavior representation learning |
0.7 | 1 | 2023 | MABe22: A Multi-Species Multi-Task Benchmark for Learned Representations of Behavior · ICML 2023 |
Computer vision › Video understanding and tracking
video representation learning |
0.7 | 1 | 2023 | MABe22: A Multi-Species Multi-Task Benchmark for Learned Representations of Behavior · ICML 2023 |
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning |
0.4 | 1 | 2020 | Robust Disentanglement of a Few Factors at a Time · NeurIPS 2020 |
Computer vision › Image recognition and object detection
object localization |
0.3 | 1 | 2025 | Probing the Mid-level Vision Capabilities of Self-Supervised Learning · CVPR 2025 |
Bioinformatics and computational biology › behavioral analysis
animal behavior analysis |
0.2 | 1 | 2023 | MABe22: A Multi-Species Multi-Task Benchmark for Learned Representations of Behavior · ICML 2023 |
Methods — techniques the papers use, named apart from their topics
self-supervised learning · 1.3pose tracking · 1.3linear probing · 0.9k-nearest neighbors · 0.9fine-tuning · 0.9controlled evaluation · 0.9benchmark protocols · 0.9model personalization · 0.8diffusion model · 0.8caption generation · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Diffusion-Based Action Recognition Generalizes to Untrained DomainsabstractHumans can recognize the same actions despite large context and viewpoint variations, such as differences between species (walking in spiders vs. horses), viewpoints (egocentric vs. third-person), and contexts (real life vs movies). Current deep learning models struggle with such generalization. We propose using features generated by a Vision Diffusion Model (VDM), aggregated via a transformer, to achieve human-like action recognition across these challenging conditions. We find that generalization is enhanced by the use of a model conditioned on earlier timesteps of the diffusion process to highlight semantic information over pixel level details in the extracted features. We experimentally explore the generalization properties of our approach in classifying actions across animal species, across different viewing angles, and different recording contexts. Our model sets a new state-of-the-art across all three generalization benchmarks, bringing machine action recognition closer to human-like robustness. Project page: vision.caltech.edu/actiondiff Code: github.com/frankyaoxiao/ActionDiff Rogério Guimarães, Frank Xiao, Pietro Perona, Markus Marks |
WACV | 4 |
| 2026 | SAVeD: Learning to Denoise Low-SNR Video for Improved Downstream PerformanceabstractLow signal-to-noise ratio (SNR) videos—such as those from underwater sonar, ultrasound, and microscopy—pose significant challenges for computer vision models, particularly when paired clean imagery for denoising is unavailable. We present Spatiotemporal Augmentations and denoising in Video for Downstream Tasks (SAVeD), a novel self-supervised method that denoises low-SNR sensor videos using only raw noisy data. By leveraging distinctions between foreground and background motion and exaggerating objects with stronger motion signal, SAVeD enhances foreground object visibility and reduces background and camera noise without requiring clean video. SAVeD has a set of architectural optimizations that lead to faster throughput, training, and inference than existing deep learning methods. We also introduce a new denoising metric, FBD, which indicates foreground-background divergence for detection datasets without requiring clean imagery. Our approach achieves state-of-the-art results for classification, detection, tracking, and counting tasks and it does so with fewer training resource requirements than existing deep-learning-based denoising methods. Project page here, Code: https://github.com/suzanne-stathatos/SAVeD. Suzanne Stathatos, Michael Hobley, Pietro Perona, Markus Marks |
WACV | 4 |
| 2025 | Probing the Mid-level Vision Capabilities of Self-Supervised LearningabstractMid-level vision capabilities — such as generic object localization and 3D geometric understanding — are not only fundamental to human vision but are also crucial for many real-world applications of computer vision. These abilities emerge with minimal supervision during the early stages of human visual development. Despite their significance, current self-supervised learning (SSL) approaches are primarily designed and evaluated for high-level recognition tasks, leaving their mid-level vision capabilities largely unexamined.In this study, we introduce a suite of benchmark protocols to systematically assess mid-level vision capabilities and present a comprehensive, controlled evaluation of 22 prominent SSL models across 8 mid-level vision tasks. Our experiments reveal a weak correlation between mid-level and high-level task performance. We also identify several SSL methods with highly imbalanced performance across mid-level and high-level capabilities, as well as some that excel in both. Additionally, we investigate key factors contributing to mid-level vision performance, such as pretraining objectives and network architectures. Our study provides a holistic and timely view of what SSL models have learned, complementing existing research that primarily focuses on high-level vision tasks. We hope our findings guide future SSL research to benchmark models not only on high-level vision tasks but on mid-level as well. Xuweiyi Chen, Markus Marks, Zezhou Cheng |
CVPR | 2 |
| 2025 | Learning Keypoints for Multi-Agent Behavior Analysis using Self-SupervisionabstractThe study of social interactions and collective behaviors through multi-agent video analysis is crucial in biology. While self-supervised keypoint discovery has emerged as a promising solution to reduce the need for manual keypoint annotations, existing methods often struggle with videos containing multiple interacting agents, especially those of the same species and color. To address this, we introduce B-KinD-multi, a novel approach that leverages pre-trained video segmentation models to guide keypoint discovery in multi-agent scenarios. This eliminates the need for time-consuming manual annotations on new experimental settings and organisms. Extensive evaluations demonstrate improved keypoint regression and downstream behavioral classification in videos of flies, mice, and rats. Furthermore, our method generalizes well to other species, including ants, bees, and humans, highlighting its potential for broad applications in automated keypoint annotation for multi-agent behavior analysis. Code available under: B-KinD-Multi Daniel Khalil, Christina Liu, Pietro Perona, Jennifer J. Sun, Markus Marks |
WACV | 5 |
| 2025 | A Closer Look at Benchmarking Self-supervised Pre-training with Image ClassificationabstractSelf-supervised learning (SSL) is a machine learning approach where the data itself provides supervision, eliminating the need for external labels. The model is forced to learn about the data's inherent structure or context by solving a pretext task. With SSL, models can learn from abundant and cheap unlabeled data, significantly reducing the cost of training models where labels are expensive or inaccessible. In Computer Vision, SSL is widely used as pre-training followed by a downstream task, such as supervised transfer, few-shot learning on smaller labeled data sets, and/or unsupervised clustering. Unfortunately, it is infeasible to evaluate SSL methods on all possible downstream tasks and objectively measure the quality of the learned representation. Instead, SSL methods are evaluated using in-domain evaluation protocols, such as fine-tuning, linear probing, and k-nearest neighbors (kNN). However, it is not well understood how well these evaluation protocols estimate the representation quality of a pre-trained model for different downstream tasks under different conditions, such as dataset, metric, and model architecture. In this work, we study how classification-based evaluation protocols for SSL correlate and how well they predict downstream performance on different dataset types. Our study includes eleven common image datasets and 26 models that were pre-trained with different SSL methods or have different model backbones. We find that in-domain linear/kNN probing protocols are, on average, the best general predictors for out-of-domain performance. We further investigate the importance of batch normalization for the various protocols and evaluate how robust correlations are for different kinds of dataset domain shifts. In addition, we challenge assumptions about the relationship between discriminative and generative self-supervised methods, finding that most of their performance differences can be explained by changes to model backbones. Supplementary Information: The online version contains supplementary material available at 10.1007/s11263-025-02402-w. Markus Marks, Manuel Knott 0001, Neehar Kondapaneni, Elijah Cole, Thijs Defraeye, Fernando Pérez-Cruz, Pietro Perona |
Int. J. Comput. Vis. | 1 |
| 2024 | Text-Image Alignment for Diffusion-Based PerceptionabstractDiffusion models are generative models with impressive text-to-image synthesis capabilities and have spurred a new wave of creative methods for classical machine learning tasks. However, the best way to harness the perceptual knowledge of these generative models for visual tasks is still an open question. Specifically, it is unclear how to use the prompting interface when applying diffusion backbones to vision tasks. We find that automatically generated captions can improve text-image alignment and significantly enhance a model's cross-attention maps, leading to better perceptual performance. Our approach improves upon the current state-of-the-art (SOTA) in diffusionbased semantic segmentation on ADE20K and the current overall SOTA for depth estimation on NYUv2. Furthermore, our method generalizes to the cross-domain setting. We use model personalization and caption modifications to align our model to the target domain and find improvements over unaligned baselines. Our crossdomain object detection model, trained on Pascal VOC, achieves SOTA results on Watercolor2K. Our cross-domain segmentation method, trained on Cityscapes, achieves SOTA results on Dark Zurich-val and Nighttime Driving. Project page: vision.caltech.edu/TADP/ Code page: github.com/damaggu/TADP Neehar Kondapaneni, Markus Marks, Manuel Knott 0001, Rogério Guimarães, Pietro Perona |
CVPR | 2 |
| 2023 | MABe22: A Multi-Species Multi-Task Benchmark for Learned Representations of BehaviorabstractWe introduce MABe22, a large-scale, multi-agent video and trajectory benchmark to assess the quality of learned behavior representations. This dataset is collected from a variety of biology experiments, and includes triplets of interacting mice (4.7 million frames video+pose tracking data, 10 million frames pose only), symbiotic beetle-ant interactions (10 million frames video data), and groups of interacting flies (4.4 million frames of pose tracking data). Accompanying these data, we introduce a panel of real-life downstream analysis tasks to assess the quality of learned representations by evaluating how well they preserve information about the experimental conditions (e.g. strain, time of day, optogenetic stimulation) and animal behavior. We test multiple state-of-the-art self-supervised video and trajectory representation learning methods to demonstrate the use of our benchmark, revealing that methods developed using human action datasets do not fully translate to animal datasets. We hope that our benchmark and dataset encourage a broader exploration of behavior representation learning methods across species and settings. Jennifer J. Sun, Markus Marks, Andrew Ulmer, Dipam Chakraborty, Brian Geuther, Edward Hayes, Heng Jia, Sebastian Oleszko, Zachary Partridge, Milan Peelman, Alice Robie, Catherine E. Schretter, Keith Sheppard, Param Uttarwar, Julian Morgan Wagner, Erik Werner, Joseph Parker, Pietro Perona, Yisong Yue, Kristin Branson, Ann Kennedy |
ICML | 2 |
| 2020 | Robust Disentanglement of a Few Factors at a Time
Benjamin Estermann, Markus Marks, Mehmet Fatih Yanik |
NeurIPS | 2 |