Frank Fundel

dblp:357/3220 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Generative modeling · 73% Efficient and distributed learning · 17% Video understanding and tracking · 5%
Computer graphics and multimedia
1 paper
Computer animation and physical simulation · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling › diffusion model › diffusion-based representation learning
diffusion model features
0.912025
CleanDIFT: Diffusion Features without Noise · CVPR 2025
Machine learning › Generative modeling
flow matching
0.912025
Diff2Flow: Training Flow Matching Models via Diffusion Model Alignment · CVPR 2025
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.912025
Diff2Flow: Training Flow Matching Models via Diffusion Model Alignment · CVPR 2025
Machine learning › Generative modeling › video generation
text-to-video synthesis
0.912025
DisMo: Disentangled Motion Representations for Open-World Motion Transfer · NeurIPS 2025
Machine learning › Generative modeling
video generation
0.912025
DisMo: Disentangled Motion Representations for Open-World Motion Transfer · NeurIPS 2025
Computer animation and physical simulation
motion transfer
0.912025
DisMo: Disentangled Motion Representations for Open-World Motion Transfer · NeurIPS 2025
Computer vision › Video understanding and tracking
action recognition
0.312025
DisMo: Disentangled Motion Representations for Open-World Motion Transfer · NeurIPS 2025
Machine learning › Generative modeling
diffusion model
0.312025
CleanDIFT: Diffusion Features without Noise · CVPR 2025
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning
0.312025
DisMo: Disentangled Motion Representations for Open-World Motion Transfer · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

video generator · 1.7image-space reconstruction · 1.7diffusion model · 1.7unsupervised fine-tuning · 0.9timestep rescaling · 0.9lightweight adapters · 0.9lightweight adapter · 0.9interpolant alignment · 0.9flow matching · 0.9feature extraction · 0.9
YearPublicationVenuePosition
2025 Diff2Flow: Training Flow Matching Models via Diffusion Model Alignment
abstract
Diffusion models have revolutionized generative tasks through high-fidelity outputs, yet flow matching (FM) offers faster inference and empirical performance gains. However, current foundation FM models are computationally prohibitive for finetuning, while diffusion models like Stable Diffusion benefit from efficient architectures and ecosystem support. This work addresses the critical challenge of efficiently transferring knowledge from pre-trained diffusion models to flow matching. We propose Diff2Flow, a novel framework that systematically bridges diffusion and FM paradigms by rescaling timesteps, aligning interpolants, and deriving FM-compatible velocity fields from diffusion predictions. This alignment enables direct and efficient FM finetuning of diffusion priors with no extra computation overhead. Our experiments demonstrate that Diff2Flow outperforms naïve FM and diffusion finetuning particularly under parameter-efficient constraints, while achieving superior or competitive performance across diverse downstream tasks compared to state-of-the-art methods. We will release our code at https://github.com/CompVis/diff2flow.
Johannes Schusterbauer, Ming Gui, Frank Fundel, Björn Ommer
CVPR3
2025 CleanDIFT: Diffusion Features without Noise
abstract
Internal features from large-scale pre-trained diffusion models have recently been established as powerful semantic descriptors for a wide range of downstream tasks. Works that use these features generally need to add noise to images before passing them through the model to obtain the semantic features, as the models do not offer the most useful features when given images with little to no noise. We show that this noise has a critical impact on the usefulness of these features that cannot be remedied by ensembling with different random noises. We address this issue by introducing a lightweight, unsupervised fine-tuning method that enables diffusion backbones to provide high-quality, noise-free semantic features. We show that these features readily outperform previous diffusion features by a wide margin in a wide variety of extraction setups and downstream tasks, offering better performance than even ensemble-based methods at a fraction of the cost.
Nick Stracke, Stefan Andreas Baumann, Kolja Bauer, Frank Fundel, Björn Ommer
CVPR4
2025 DisMo: Disentangled Motion Representations for Open-World Motion Transfer
abstract
Recent advances in text-to-video (T2V) and image-to-video (I2V) models, have enabled the creation of visually compelling and dynamic videos from simple textual descriptions or initial frames. However, these models often fail to provide an explicit representation of motion separate from content, limiting their applicability for content creators. To address this gap, we propose DisMo, a novel paradigm for learning abstract motion representations directly from raw video data via an image-space reconstruction objective. Our representation is generic and independent of static information such as appearance, object identity, or pose. This enables open-world motion transfer, allowing motion to be transferred across semantically unrelated entities without requiring object correspondences, even between vastly different categories. Unlike prior methods, which trade off motion fidelity and prompt adherence, are overfitting to source structure or drifting from the described action, our approach disentangles motion semantics from appearance, enabling accurate transfer and faithful conditioning. Furthermore, our motion representation can be combined with any existing video generator via lightweight adapters, allowing us to effortlessly benefit from future advancements in video models. We demonstrate the effectiveness of our method through a diverse set of motion transfer tasks. Finally, we show that the learned representations are well-suited for downstream motion understanding tasks, consistently outperforming state-of-the-art video representation models such as V-JEPA in zero-shot action classification on benchmarks including Something-Something v2 and Jester. Project page: https://compvis.github.io/DisMo
Thomas Ressler-Antal, Frank Fundel, Malek Ben Alaya, Stefan Andreas Baumann, Felix Krause 0002, Ming Gui, Björn Ommer
NeurIPS2
2025 Distillation of Diffusion Features for Semantic Correspondence
abstract
Semantic correspondence, the task of determining relationships between different parts of images, underpins various applications including 3D reconstruction, image-to-image translation, object tracking, and visual place recognition. Recent studies have begun to explore representations learned in large generative image models for semantic correspondence, demonstrating promising results. Building on this progress, current state-of-the-art methods rely on combining multiple large models, resulting in high computational demands and reduced efficiency. In this work, we address this challenge by proposing a more computationally efficient approach. We propose a novel knowledge distillation technique to overcome the problem of reduced efficiency. We show how to use two large vision foundation models and distill the capabilities of these complementary models into one smaller model that maintains high accuracy at reduced computational cost. Furthermore, we demonstrate that by incorporating 3D data, we are able to further improve performance, without the need for human-annotated correspondences. Overall, our empirical results demonstrate that our distilled model with 3D data augmentation achieves performance superior to current state-of-the-art methods while significantly reducing computational load and enhancing practicality for real-world applications, such as semantic video correspondence. Our code and weights are publicly available on our project page.
Frank Fundel, Johannes Schusterbauer, Vincent Tao Hu, Björn Ommer
WACV1