Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Olga Zatsarynna

dblp:255/6446 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Video understanding and tracking · 45% Generative modeling · 35% Deep learning architectures and training · 17%

Topics — the 8 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
2.532025
MANTA: Diffusion Mamba for Efficient and Effective Stochastic Long-Term Dense Action Anticipation · CVPR 2025
SyncVP: Joint Diffusion for Synchronous Multi-Modal Video Prediction · CVPR 2025
Gated Temporal Diffusion for Stochastic Long-Term Dense Anticipation · ECCV (55) 2024
Computer vision › Video understanding and tracking
action anticipation
1.722025
MixANT: Observation-Dependent Memory Propagation for Stochastic Dense Action Anticipation · ICCV 2025
MANTA: Diffusion Mamba for Efficient and Effective Stochastic Long-Term Dense Action Anticipation · CVPR 2025
Machine learning › Deep learning architectures and training
state space model
1.122025
MixANT: Observation-Dependent Memory Propagation for Stochastic Dense Action Anticipation · ICCV 2025
MANTA: Diffusion Mamba for Efficient and Effective Stochastic Long-Term Dense Action Anticipation · CVPR 2025
Computer vision › Video understanding and tracking
action segmentation
1.012026
Towards Generalizing Temporal Action Segmentation to Unseen Views · Int. J. Comput. Vis. 2026
Machine learning › Generative modeling › diffusion model
multimodal diffusion model
0.912025
SyncVP: Joint Diffusion for Synchronous Multi-Modal Video Prediction · CVPR 2025
Computer vision › Video understanding and tracking
video prediction
0.912025
SyncVP: Joint Diffusion for Synchronous Multi-Modal Video Prediction · CVPR 2025
Machine learning › Deep learning architectures and training
mixture of experts
0.312025
MixANT: Observation-Dependent Memory Propagation for Stochastic Dense Action Anticipation · ICCV 2025
Machine learning › Deep learning architectures and training
sequence modeling
0.312025
MANTA: Diffusion Mamba for Efficient and Effective Stochastic Long-Term Dense Action Anticipation · CVPR 2025

Methods — techniques the papers use, named apart from their topics

diffusion · 1.6shared representation learning · 1.0sequence loss · 1.0action loss · 1.0stochastic prediction · 0.9spatio-temporal cross-attention · 0.9mixture of experts · 0.9mamba · 0.9forget-gate mechanism · 0.9diffusion model · 0.9
YearPublicationVenuePosition
2026 Looking into the unknown: Exploring Action Discovery for segmentation of known and unknown actions
abstract
We introduce Action Discovery, a novel task that addresses the challenge of discovering actions in long, untrimmed videos where only a subset of the present actions have been annotated. The goalis thus to discover new actions in the video segments that have not been annotated or annotated by a generic background class. This scenario is particularly relevant in domains like neuroscience, where well-defined behaviors (e.g., walking, eating) coexist with subtle or infrequent actions that are often overlooked, as well as in applications where datasets are inherently partially annotateddue to ambiguous or missing labels. To address this problem, we propose a two-step approach that leverages the known annotations to guide both the temporal and semantic granularity of unknownaction segments. First, we introduce the Granularity-Guided Segmentation Module (GGSM), which identifies temporal intervals for both known and unknown actions by mimicking the granularity ofannotated actions. Second, we propose the Unknown Action Segment Assignment (UASA), which identifies semantically meaningful classes within the unknown actions, based on learned embedding similarities. We systematically explore the proposed setting of Action Discovery on three challenging datasets - Breakfast, 50Salads, and Desktop Assembly - demonstrating that our method considerablyimproves upon existing baselines.
Federico Spurio, Emad Bahrami, Olga Zatsarynna, Yazan Abu Farha, Gianpiero Francesca, Juergen Gall
Comput. Vis. Image Underst.3
2026 Towards Generalizing Temporal Action Segmentation to Unseen Views
abstract
Abstract While there has been substantial progress in temporal action segmentation, the challenge to generalize to unseen views remains unaddressed. Hence, we define a protocol for unseen view action segmentation where camera views for evaluating the model are unavailable during training. This includes changing from top-frontal views to a side view or even more challenging from exocentric to egocentric views. Furthermore, we present an approach for temporal action segmentation that tackles this challenge. Our approach leverages a shared representation at both the sequence and segment levels to reduce the impact of view differences during training. We achieve this by introducing a sequence loss and an action loss, which together facilitate consistent video and action representations across different views. The evaluation on the Assembly101, IkeaASM, and EgoExoLearn datasets demonstrate significant improvements, with a $$12.8\%$$ 12.8 % increase in F1@50 for unseen exocentric views and a substantial $$54\%$$ 54 % improvement for unseen egocentric views.
Emad Bahrami, Olga Zatsarynna, Gianpiero Francesca, Juergen Gall
Int. J. Comput. Vis.2
2025 SyncVP: Joint Diffusion for Synchronous Multi-Modal Video Prediction
abstract
Predicting future video frames is essential for decision-making systems, yet RGB frames alone often lack the information needed to fully capture the underlying complexities of the real world. To address this limitation, we propose a multi-modal framework for Synchronous Video Prediction (SyncVP) that incorporates complementary data modalities, enhancing the richness and accuracy of future predictions. SyncVP builds on pre-trained modality-specific diffusion models and introduces an efficient spatio-temporal cross-attention module to enable effective information sharing across modalities. We evaluate SyncVP on standard benchmark datasets, such as Cityscapes and BAIR, using depth as an additional modality. We furthermore demonstrate its generalization to other modalities on SYNTHIA with semantic information and ERA5-Land with climate data. Notably, SyncVP achieves state-of-the-art performance, even in scenarios where only one modality is present, demonstrating its robustness and potential for a wide range of applications.
Enrico Pallotta, Sina Mokhtarzadeh Azar, Olga Zatsarynna, Juergen Gall
CVPR4
2025 MANTA: Diffusion Mamba for Efficient and Effective Stochastic Long-Term Dense Action Anticipation
abstract
Long-term dense action anticipation is very challenging since it requires predicting actions and their durations several minutes into the future based on provided video observations. To model the uncertainty of future outcomes, stochastic models predict several potential future action sequences for the same observation. Recent work has further proposed to incorporate uncertainty modelling for observed frames by simultaneously predicting per-frame past and future actions in a unified manner. While such joint modelling of actions is beneficial, it requires long-range temporal capabilities to connect events across distant past and future time points. However, the previous work struggles to achieve such a long-range understanding due to its limited and/or sparse receptive field. To alleviate this issue, we propose a novel MANTA (MAmbafor ANTicipation) network. Our model enables effective long-term temporal modelling even for very long sequences while maintaining linear complexity in sequence length. We demonstrate that our approach achieves state-of-the-art results on three datasets—Breakfast, 50Salads, and Assembly 101—while also significantly improving computational and memory efficiency. Our code is available at https://github.com/olga-zats/DlFFMANTA.
Olga Zatsarynna, Emad Bahrami, Yazan Abu Farha, Gianpiero Francesca, Juergen Gall
CVPR1
2025 MixANT: Observation-Dependent Memory Propagation for Stochastic Dense Action Anticipation
abstract
We present MixANT, a novel architecture for stochastic long-term dense anticipation of human activities. While recent State Space Models (SSMs) like Mamba have shown promise through input-dependent selectivity on three key parameters, the critical forget-gate ($\textbf{A}$ matrix) controlling temporal memory remains static. We address this limitation by introducing a mixture of experts approach that dynamically selects contextually relevant $\textbf{A}$ matrices based on input features, enhancing representational capacity without sacrificing computational efficiency. Extensive experiments on the 50Salads, Breakfast, and Assembly101 datasets demonstrate that MixANT consistently outperforms state-of-the-art methods across all evaluation settings. Our results highlight the importance of input-dependent forget-gate mechanisms for reliable prediction of human behavior in diverse real-world scenarios.
Syed Talal Wasim, Hamid Suleman, Olga Zatsarynna, Muzammal Naseer, Juergen Gall
ICCV3
2024 Gated Temporal Diffusion for Stochastic Long-Term Dense Anticipation
Olga Zatsarynna, Emad Bahrami, Yazan Abu Farha, Gianpiero Francesca, Juergen Gall
ECCV (55)1
2023 Action Anticipation with Goal Consistency
abstract
In this paper, we address the problem of short-term action anticipation, i.e., we want to predict an upcoming action one second before it happens. We propose to harness high-level intent information to anticipate actions that will take place in the future. To this end, we incorporate an additional goal prediction branch into our model and propose a consistency loss function that encourages the anticipated actions to conform to the high-level goal pursued in the video. In our experiments, we show the effectiveness of the proposed approach and demonstrate that our method achieves state-of-the-art results on two large-scale datasets: Assembly101 and COIN. The code is available at https://github.com/olga-zats/goal_consistency.
Olga Zatsarynna, Juergen Gall
ICIP1