Nick Stracke

dblp:364/1346 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Generative modeling · 78% Video understanding and tracking · 9% Segmentation and scene understanding · 9%
Computer graphics and multimedia
1 paper
Computer animation and physical simulation · 100%

Topics — the 9 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.832025
CTRLorALTer: Conditional LoRAdapter for Efficient 0-Shot Control and Altering of T2I Models · ECCV (88) 2024
FMBoost: Boosting Latent Diffusion with Flow Matching · ECCV (61) 2024
CleanDIFT: Diffusion Features without Noise · CVPR 2025
Machine learning › Generative modeling › diffusion model › diffusion-based representation learning
diffusion model features
0.912025
CleanDIFT: Diffusion Features without Noise · CVPR 2025
Computer vision › Video understanding and tracking › motion analysis
scene dynamics
0.912025
What If: Understanding Motion Through Sparse Interactions · ICCV 2025
Machine learning › Generative modeling › diffusion model › text-to-image generation
text-to-image diffusion model
0.912025
Continuous, Subject-Specific Attribute Control in T2I Models by Identifying Semantic Directions · CVPR 2025
Machine learning › Generative modeling › diffusion model
controllable generation
0.812024
CTRLorALTer: Conditional LoRAdapter for Efficient 0-Shot Control and Altering of T2I Models · ECCV (88) 2024
Machine learning › Generative modeling
flow matching
0.812024
FMBoost: Boosting Latent Diffusion with Flow Matching · ECCV (61) 2024
Machine learning › Generative modeling › diffusion model
image editing
0.812024
CTRLorALTer: Conditional LoRAdapter for Efficient 0-Shot Control and Altering of T2I Models · ECCV (88) 2024
Machine learning › Generative modeling › diffusion model
latent diffusion model
0.812024
FMBoost: Boosting Latent Diffusion with Flow Matching · ECCV (61) 2024
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.812024
CTRLorALTer: Conditional LoRAdapter for Efficient 0-Shot Control and Altering of T2I Models · ECCV (88) 2024

Methods — techniques the papers use, named apart from their topics

flow matching · 2.5transformer · 1.7diffusion model · 1.7unsupervised fine-tuning · 0.9optimization-free direction identification · 0.9learning-based semantic direction identification · 0.9feature extraction · 0.9latent diffusion · 0.8adapter · 0.8LoRA · 0.8
YearPublicationVenuePosition
2025 Continuous, Subject-Specific Attribute Control in T2I Models by Identifying Semantic Directions
abstract
Recent advances in text-to-image (T2I) diffusion models have significantly improved the quality of generated images. However, providing efficient control over individual subjects, particularly the attributes characterizing them, remains a key challenge. While existing methods have introduced mechanisms to modulate attribute expression, they typically provide either detailed, object-specific localization of such a modification or full-scale fine-grained, nuanced control of attributes. No current approach offers both simultaneously, resulting in a gap when trying to achieve precise continuous and subject-specific attribute modulation in image generation. In this work, we demonstrate that token-level directions exist within commonly used CLIP text embeddings that enable fine-grained, subject-specific control of high-level attributes in T2I models. We introduce two methods to identify these directions: a simple, optimization-free technique and a learning-based approach that utilizes the T2I model to characterize semantic concepts more specifically. Our methods allow the augmentation of the prompt text input, enabling fine-grained control over multiple attributes of individual subjects simultaneously, without requiring any modifications to the diffusion model itself. This approach offers a unified solution that fills the gap between global and localized control, providing competitive flexibility and precision in text-guided image generation.
Stefan Andreas Baumann, Felix Krause 0002, Michael Neumayr, Nick Stracke, Melvin Sevi, Vincent Tao Hu, Björn Ommer
CVPR4
2025 CleanDIFT: Diffusion Features without Noise
abstract
Internal features from large-scale pre-trained diffusion models have recently been established as powerful semantic descriptors for a wide range of downstream tasks. Works that use these features generally need to add noise to images before passing them through the model to obtain the semantic features, as the models do not offer the most useful features when given images with little to no noise. We show that this noise has a critical impact on the usefulness of these features that cannot be remedied by ensembling with different random noises. We address this issue by introducing a lightweight, unsupervised fine-tuning method that enables diffusion backbones to provide high-quality, noise-free semantic features. We show that these features readily outperform previous diffusion features by a wide margin in a wide variety of extraction setups and downstream tasks, offering better performance than even ensemble-based methods at a fraction of the cost.
Nick Stracke, Stefan Andreas Baumann, Kolja Bauer, Frank Fundel, Björn Ommer
CVPR1
2025 What If: Understanding Motion Through Sparse Interactions
abstract
Understanding the dynamics of a physical scene involves reasoning about the diverse ways it can potentially change, especially as a result of local interactions. We present the Flow Poke Transformer (FPT), a novel framework for directly predicting the distribution of local motion, conditioned on sparse interactions termed "pokes". Unlike traditional methods that typically only enable dense sampling of a single realization of scene dynamics, FPT provides an interpretable directly accessible representation of multi-modal scene motion, its dependency on physical interactions and the inherent uncertainties of scene dynamics. We also evaluate our model on several downstream tasks to enable comparisons with prior methods and highlight the flexibility of our approach. On dense face motion generation, our generic pre-trained model surpasses specialized baselines. FPT can be fine-tuned in strongly out-of-distribution tasks such as synthetic datasets to enable significant improvements over in-domain methods in articulated object motion estimation. Additionally, predicting explicit motion distributions directly enables our method to achieve competitive performance on tasks like moving part segmentation from pokes which further demonstrates the versatility of our FPT. Code and models are publicly available at https://compvis.github.io/flow-poke-transformer.
Stefan Andreas Baumann, Nick Stracke, Timy Phan, Björn Ommer
ICCV2
2024 FMBoost: Boosting Latent Diffusion with Flow Matching
Johannes Schusterbauer, Ming Gui, Pingchuan Ma 0006, Nick Stracke, Stefan Andreas Baumann, Vincent Tao Hu, Björn Ommer
ECCV (61)4
2024 CTRLorALTer: Conditional LoRAdapter for Efficient 0-Shot Control and Altering of T2I Models
Nick Stracke, Stefan Andreas Baumann, Joshua Susskind, Miguel Ángel Bautista 0001, Björn Ommer
ECCV (88)1