VLDB 2026 Research / reviewers in the wild / expert
Olga Zatsarynna
dblp:255/6446
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Video understanding and tracking · 45% Generative modeling · 35% Deep learning architectures and training · 17% |
Topics — the 8 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
2.5 | 3 | 2025 | MANTA: Diffusion Mamba for Efficient and Effective Stochastic Long-Term Dense Action Anticipation · CVPR 2025 SyncVP: Joint Diffusion for Synchronous Multi-Modal Video Prediction · CVPR 2025 Gated Temporal Diffusion for Stochastic Long-Term Dense Anticipation · ECCV (55) 2024 |
Computer vision › Video understanding and tracking
action anticipation |
1.7 | 2 | 2025 | MixANT: Observation-Dependent Memory Propagation for Stochastic Dense Action Anticipation · ICCV 2025 MANTA: Diffusion Mamba for Efficient and Effective Stochastic Long-Term Dense Action Anticipation · CVPR 2025 |
Machine learning › Deep learning architectures and training
state space model |
1.1 | 2 | 2025 | MixANT: Observation-Dependent Memory Propagation for Stochastic Dense Action Anticipation · ICCV 2025 MANTA: Diffusion Mamba for Efficient and Effective Stochastic Long-Term Dense Action Anticipation · CVPR 2025 |
Computer vision › Video understanding and tracking
action segmentation |
1.0 | 1 | 2026 | Towards Generalizing Temporal Action Segmentation to Unseen Views · Int. J. Comput. Vis. 2026 |
Machine learning › Generative modeling › diffusion model
multimodal diffusion model |
0.9 | 1 | 2025 | SyncVP: Joint Diffusion for Synchronous Multi-Modal Video Prediction · CVPR 2025 |
Computer vision › Video understanding and tracking
video prediction |
0.9 | 1 | 2025 | SyncVP: Joint Diffusion for Synchronous Multi-Modal Video Prediction · CVPR 2025 |
Machine learning › Deep learning architectures and training
mixture of experts |
0.3 | 1 | 2025 | MixANT: Observation-Dependent Memory Propagation for Stochastic Dense Action Anticipation · ICCV 2025 |
Machine learning › Deep learning architectures and training
sequence modeling |
0.3 | 1 | 2025 | MANTA: Diffusion Mamba for Efficient and Effective Stochastic Long-Term Dense Action Anticipation · CVPR 2025 |
Methods — techniques the papers use, named apart from their topics
diffusion · 1.6shared representation learning · 1.0sequence loss · 1.0action loss · 1.0stochastic prediction · 0.9spatio-temporal cross-attention · 0.9mixture of experts · 0.9mamba · 0.9forget-gate mechanism · 0.9diffusion model · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Looking into the unknown: Exploring Action Discovery for segmentation of known and unknown actionsabstractWe introduce Action Discovery, a novel task that addresses the challenge of discovering actions in long, untrimmed videos where only a subset of the present actions have been annotated. The goalis thus to discover new actions in the video segments that have not been annotated or annotated by a generic background class. This scenario is particularly relevant in domains like neuroscience, where well-defined behaviors (e.g., walking, eating) coexist with subtle or infrequent actions that are often overlooked, as well as in applications where datasets are inherently partially annotateddue to ambiguous or missing labels. To address this problem, we propose a two-step approach that leverages the known annotations to guide both the temporal and semantic granularity of unknownaction segments. First, we introduce the Granularity-Guided Segmentation Module (GGSM), which identifies temporal intervals for both known and unknown actions by mimicking the granularity ofannotated actions. Second, we propose the Unknown Action Segment Assignment (UASA), which identifies semantically meaningful classes within the unknown actions, based on learned embedding similarities. We systematically explore the proposed setting of Action Discovery on three challenging datasets - Breakfast, 50Salads, and Desktop Assembly - demonstrating that our method considerablyimproves upon existing baselines. Federico Spurio, Emad Bahrami, Olga Zatsarynna, Yazan Abu Farha, Gianpiero Francesca, Juergen Gall |
Comput. Vis. Image Underst. | 3 |
| 2026 | Towards Generalizing Temporal Action Segmentation to Unseen ViewsabstractAbstract While there has been substantial progress in temporal action segmentation, the challenge to generalize to unseen views remains unaddressed. Hence, we define a protocol for unseen view action segmentation where camera views for evaluating the model are unavailable during training. This includes changing from top-frontal views to a side view or even more challenging from exocentric to egocentric views. Furthermore, we present an approach for temporal action segmentation that tackles this challenge. Our approach leverages a shared representation at both the sequence and segment levels to reduce the impact of view differences during training. We achieve this by introducing a sequence loss and an action loss, which together facilitate consistent video and action representations across different views. The evaluation on the Assembly101, IkeaASM, and EgoExoLearn datasets demonstrate significant improvements, with a $$12.8\%$$ 12.8 % increase in F1@50 for unseen exocentric views and a substantial $$54\%$$ 54 % improvement for unseen egocentric views. Emad Bahrami, Olga Zatsarynna, Gianpiero Francesca, Juergen Gall |
Int. J. Comput. Vis. | 2 |
| 2025 | SyncVP: Joint Diffusion for Synchronous Multi-Modal Video PredictionabstractPredicting future video frames is essential for decision-making systems, yet RGB frames alone often lack the information needed to fully capture the underlying complexities of the real world. To address this limitation, we propose a multi-modal framework for Synchronous Video Prediction (SyncVP) that incorporates complementary data modalities, enhancing the richness and accuracy of future predictions. SyncVP builds on pre-trained modality-specific diffusion models and introduces an efficient spatio-temporal cross-attention module to enable effective information sharing across modalities. We evaluate SyncVP on standard benchmark datasets, such as Cityscapes and BAIR, using depth as an additional modality. We furthermore demonstrate its generalization to other modalities on SYNTHIA with semantic information and ERA5-Land with climate data. Notably, SyncVP achieves state-of-the-art performance, even in scenarios where only one modality is present, demonstrating its robustness and potential for a wide range of applications. Enrico Pallotta, Sina Mokhtarzadeh Azar, Olga Zatsarynna, Juergen Gall |
CVPR | 4 |
| 2025 | MANTA: Diffusion Mamba for Efficient and Effective Stochastic Long-Term Dense Action AnticipationabstractLong-term dense action anticipation is very challenging since it requires predicting actions and their durations several minutes into the future based on provided video observations. To model the uncertainty of future outcomes, stochastic models predict several potential future action sequences for the same observation. Recent work has further proposed to incorporate uncertainty modelling for observed frames by simultaneously predicting per-frame past and future actions in a unified manner. While such joint modelling of actions is beneficial, it requires long-range temporal capabilities to connect events across distant past and future time points. However, the previous work struggles to achieve such a long-range understanding due to its limited and/or sparse receptive field. To alleviate this issue, we propose a novel MANTA (MAmbafor ANTicipation) network. Our model enables effective long-term temporal modelling even for very long sequences while maintaining linear complexity in sequence length. We demonstrate that our approach achieves state-of-the-art results on three datasets—Breakfast, 50Salads, and Assembly 101—while also significantly improving computational and memory efficiency. Our code is available at https://github.com/olga-zats/DlFFMANTA. Olga Zatsarynna, Emad Bahrami, Yazan Abu Farha, Gianpiero Francesca, Juergen Gall |
CVPR | 1 |
| 2025 | MixANT: Observation-Dependent Memory Propagation for Stochastic Dense Action AnticipationabstractWe present MixANT, a novel architecture for stochastic long-term dense anticipation of human activities. While recent State Space Models (SSMs) like Mamba have shown promise through input-dependent selectivity on three key parameters, the critical forget-gate ($\textbf{A}$ matrix) controlling temporal memory remains static. We address this limitation by introducing a mixture of experts approach that dynamically selects contextually relevant $\textbf{A}$ matrices based on input features, enhancing representational capacity without sacrificing computational efficiency. Extensive experiments on the 50Salads, Breakfast, and Assembly101 datasets demonstrate that MixANT consistently outperforms state-of-the-art methods across all evaluation settings. Our results highlight the importance of input-dependent forget-gate mechanisms for reliable prediction of human behavior in diverse real-world scenarios. Syed Talal Wasim, Hamid Suleman, Olga Zatsarynna, Muzammal Naseer, Juergen Gall |
ICCV | 3 |
| 2024 | Gated Temporal Diffusion for Stochastic Long-Term Dense Anticipation
Olga Zatsarynna, Emad Bahrami, Yazan Abu Farha, Gianpiero Francesca, Juergen Gall |
ECCV (55) | 1 |
| 2023 | Action Anticipation with Goal ConsistencyabstractIn this paper, we address the problem of short-term action anticipation, i.e., we want to predict an upcoming action one second before it happens. We propose to harness high-level intent information to anticipate actions that will take place in the future. To this end, we incorporate an additional goal prediction branch into our model and propose a consistency loss function that encourages the anticipated actions to conform to the high-level goal pursued in the video. In our experiments, we show the effectiveness of the proposed approach and demonstrate that our method achieves state-of-the-art results on two large-scale datasets: Assembly101 and COIN. The code is available at https://github.com/olga-zats/goal_consistency. Olga Zatsarynna, Juergen Gall |
ICIP | 1 |