VLDB 2026 Research / reviewers in the wild / expert
Michael Finkelson
dblp:405/3115
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
1 paper |
Audio and music processing · 100% | |
| Artificial intelligence
1 paper |
Generative modeling · 100% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling › audio generation
text-to-audio generation |
0.9 | 1 | 2025 | CAFA: A Controllable Automatic Foley Artist · ICCV 2025 |
Audio and music processing › sound synthesis › video-to-audio generation
foley sound generation |
0.9 | 1 | 2025 | CAFA: A Controllable Automatic Foley Artist · ICCV 2025 |
Audio and music processing › sound synthesis
video-to-audio generation |
0.9 | 1 | 2025 | CAFA: A Controllable Automatic Foley Artist · ICCV 2025 |
Methods — techniques the papers use, named apart from their topics
text-to-audio model · 1.7modality adapter · 1.7diffusion model · 1.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CAFA: A Controllable Automatic Foley ArtistabstractFoley is a key element in video production, refers to the process of adding an audio signal to a silent video while ensuring semantic and temporal alignment. In recent years, the rise of personalized content creation and advancements in automatic video-to-audio models have increased the demand for greater user control in the process. One possible approach is to incorporate text to guide audio generation. While supported by existing methods, challenges remain in ensuring compatibility between modalities, particularly when the text introduces additional information or contradicts the sounds naturally inferred from the visuals. In this work, we introduce CAFA (Controllable Automatic Foley Artist) a video-and-text-to-audio model that generates semantically and temporally aligned audio for a given video, guided by text input. CAFA is built upon a text-to-audio model and integrates video information through a modality adapter mechanism. By incorporating text, users can refine semantic details and introduce creative variations, guiding the audio synthesis beyond the expected video contextual cues. Experiments show that besides its superior quality in terms of semantic alignment and audio-visual synchronization the proposed method enable high textual controllability as demonstrated in subjective and objective evaluations. Roi Benita, Michael Finkelson, Tavi Halperin, Gleb Sterkin, Yossi Adi |
ICCV | 2 |