Pedro Vélez

dblp:395/4561 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Generative modeling · 56% Video understanding and tracking · 44%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling › diffusion model › diffusion-based representation learning
diffusion model features
0.912025
From Image to Video: An Empirical Study of Diffusion Representations · ICCV 2025
Computer vision › Video understanding and tracking
video representation learning
0.912025
From Image to Video: An Empirical Study of Diffusion Representations · ICCV 2025
Machine learning › Generative modeling
diffusion model
0.312025
From Image to Video: An Empirical Study of Diffusion Representations · ICCV 2025

Methods — techniques the papers use, named apart from their topics

representation probing · 0.9diffusion model · 0.9
YearPublicationVenuePosition
2025 From Image to Video: An Empirical Study of Diffusion Representations
abstract
Diffusion models have revolutionized generative modeling, enabling unprecedented realism in image and video synthesis. This success has sparked interest in leveraging their representations for visual understanding tasks. While recent works have explored this potential for image generation, the visual understanding capabilities of video diffusion models remain largely uncharted. To address this gap, we systematically compare the same model architecture trained for video versus image generation, analyzing the performance of their latent representations on various downstream tasks including image classification, action recognition, depth estimation, and tracking. Results show that video diffusion models consistently outperform their image counterparts, though we find a striking range in the extent of this superiority. We further analyze features extracted from different layers and with varying noise levels, as well as the effect of model size and training budget on representation and generation quality. This work marks the first direct comparison of video and image diffusion objectives for visual understanding, offering insights into the role of temporal information in representation learning.
Pedro Vélez, Luisa F. Polanía, Yi Yang 0007, Rishabh Kabra, Anurag Arnab, Mehdi S. M. Sajjadi
ICCV1