VLDB 2026 Research / reviewers in the wild / expert
Enrico Pallotta
dblp:316/6320
· DBLP profile ↗
1ranked-venue papers
1as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Generative modeling · 67% Video understanding and tracking · 33% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | SyncVP: Joint Diffusion for Synchronous Multi-Modal Video Prediction · CVPR 2025 |
Machine learning › Generative modeling › diffusion model
multimodal diffusion model |
0.9 | 1 | 2025 | SyncVP: Joint Diffusion for Synchronous Multi-Modal Video Prediction · CVPR 2025 |
Computer vision › Video understanding and tracking
video prediction |
0.9 | 1 | 2025 | SyncVP: Joint Diffusion for Synchronous Multi-Modal Video Prediction · CVPR 2025 |
Methods — techniques the papers use, named apart from their topics
spatio-temporal cross-attention · 0.9diffusion model · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SyncVP: Joint Diffusion for Synchronous Multi-Modal Video PredictionabstractPredicting future video frames is essential for decision-making systems, yet RGB frames alone often lack the information needed to fully capture the underlying complexities of the real world. To address this limitation, we propose a multi-modal framework for Synchronous Video Prediction (SyncVP) that incorporates complementary data modalities, enhancing the richness and accuracy of future predictions. SyncVP builds on pre-trained modality-specific diffusion models and introduces an efficient spatio-temporal cross-attention module to enable effective information sharing across modalities. We evaluate SyncVP on standard benchmark datasets, such as Cityscapes and BAIR, using depth as an additional modality. We furthermore demonstrate its generalization to other modalities on SYNTHIA with semantic information and ERA5-Land with climate data. Notably, SyncVP achieves state-of-the-art performance, even in scenarios where only one modality is present, demonstrating its robustness and potential for a wide range of applications. Enrico Pallotta, Sina Mokhtarzadeh Azar, Olga Zatsarynna, Juergen Gall |
CVPR | 1 |